Apodex 1.1 把任务状态放在模型消息历史之外,将可执行环境本身视为扩展面,而非给智能体堆更多工具。其 Agent Team 同时是训练好的通用模型与更大执行栈,在 GDPVal、FrontierFinance、FrontierScience-Research 上分别提升 9.3、5.6、8.3 分。
A context window is a terrible place to store the state of an 80-minute agent run.
One of the better ideas in Apodex 1.1 is that the task state sits outside the model's message history.
So @Apodex_AI is treating executable environments themselves as a scaling surface.
That is a much bigger idea than adding more tools to an agent.
So Apodex 1.1 Agent Team is simultaneously a trained general-purpose model and the larger execution stack used to turn that model into a long-running worker.
They reported gains of 9.3 points on GDPVal, 5.6 on FrontierFinance, and 8.3 on FrontierScience-Research.
Their training environments are actual file, search, and code worlds with state transitions, tool budgets, failure conditions, and task-level verifiers.
The model has to operate inside them, change the workspace, recover when something breaks, and eventually satisfy a delivery contract.
So the training distribution contains trajectories of work, not just prompts paired with good answers.
That distinction is so important for agents. You can train a model to know how a tool works and still have it fail the moment the authoritative file changes, a command errors halfway through, or two artifacts become inconsistent.
Apodex's argument is basically that the next scaling surface is the world the model learns to work inside.
来源:@rohanpaul_ai · x.com