跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-26AI 评分39
AI 导读

Apodex 1.1 把任务状态放在模型消息历史之外,将可执行环境本身视为扩展面,而非给智能体堆更多工具。其 Agent Team 同时是训练好的通用模型与更大执行栈,在 GDPVal、FrontierFinance、FrontierScience-Research 上分别提升 9.3、5.6、8.3 分。

正文

A context window is a terrible place to store the state of an 80-minute agent run.

One of the better ideas in Apodex 1.1 is that the task state sits outside the model's message history.

So @Apodex_AI is treating executable environments themselves as a scaling surface.

That is a much bigger idea than adding more tools to an agent.

So Apodex 1.1 Agent Team is simultaneously a trained general-purpose model and the larger execution stack used to turn that model into a long-running worker.

They reported gains of 9.3 points on GDPVal, 5.6 on FrontierFinance, and 8.3 on FrontierScience-Research.

Their training environments are actual file, search, and code worlds with state transitions, tool budgets, failure conditions, and task-level verifiers.

The model has to operate inside them, change the workspace, recover when something breaks, and eventually satisfy a delivery contract.

So the training distribution contains trajectories of work, not just prompts paired with good answers.

That distinction is so important for agents. You can train a model to know how a tool works and still have it fail the moment the authoritative file changes, a command errors halfway through, or two artifacts become inconsistent.

Apodex's argument is basically that the next scaling surface is the world the model learns to work inside.

来源:@rohanpaul_ai · x.com