跳到正文
@SemiAnalysis_· @SemiAnalysis_ · X·· 28 天前AI 评分29
AI 导读

这怎么可能?Terminal Bench 2.1 中的所有任务都是完全公开的。虽然 Meta 和 Google 绝不会直接在任务上训练,但他们绝对会从 RL 环境初创公司购买数据,这些数据被设计得尽可能贴近 TB 2.1 任务。最终效果是一样的。你通常会期望 TB 2.1 性能的提升能泛化到其他智能体任务上,但 Gemini 和 Muse 甚至无法泛化到 TB 4.0。(2/5)

正文

How is this possible? All of the tasks in Terminal Bench 2.1 are fully public. Though Meta and Google would never train on the tasks directly, they absolutely will buy data from RL env startups that’s designed to mimic TB 2.1 tasks as closely as possible. The net effect is the same. You’d typically expect improved TB 2.1 performance to generalize to other agentic tasks, but Gemini and Muse don't even generalize to TB 4.0. (2/5)

来源:@SemiAnalysis_ · x.com