swyx 评论 Cognition 发布首个评测,其私有企业评测时长上限达 100 小时,而 METR 约 16 小时。METR 数据集为 7 名技术人员 34 场会话、rlog 0.83;Cognition 数据集为 126 名用户 258 场会话、留出集 rlog 0.74。
Finally! the first eval ship from cog!!!!!!!!!! 👼🏼
To contextualize: @METR_Evals cap out at ~16 hours.
Cog has private enterprise evals up to 100hrs, and is confident enough to put a financial guarantee on it 🤯
METR dataset: ML eng, GPU kernels, cybersecurity
> "METR (2026) used a combination of GPT-4o and GPT-5 to estimate the human-equivalent times from compressed Claude Code transcripts. These transcripts were collected from 7 METR technical staff on 34 sessions labeled on human ground truth". rlog of 0.83
Cog dataset: real life java/typescript/python/c# feature dev, bugfixes, migrations
> "We collected a ground-truth dataset by asking Devin users to review recent representative sessions, and estimate how long each completed session would have taken without Devin. Our dataset consists of 258 sessions from 126 users across a diverse set of enterprise customers." rlog of 0.74 on held out set
this is pioneering real world evals work and part 1 of a broader frontier code evals drop that I'm really looking forward to writing up. huge kudos to @annarmitchell and @ryanbai1412 for leading the unglamorous last mile data collection!!
AI should earn its keep. Introducing the AI Productivity Guarantee. If Devin delivers less engineering value than you’re paying for, Cognition will fund your usage until it does, up to $10 million. It’s time for the AI industry to stop maximizing tokens and start maximizing productive output.在 X 查看被引用的帖子
来源:@swyx · x.com