AI 导读
评估单位是完整的求解器。 这意味着模型加上它的工具、记忆、执行环境和控制循环。 只对基础模型打分,会遗漏智能体在实践中真正失败的地方。https://t.co/Dui8MMqZZF
正文
The evaluation unit is the complete solver.
That means the model plus its tools, memory, execution environment, and control loop.
Scoring the base model alone misses where agents actually fail in practice. https://t.co/Dui8MMqZZF
来源:@omarsar0 · x.com