AI 导读
之前 @cognition_labs 的节目…… 终于!cog 的第一个 eval 发布了!!!!!!👼🏼 背景补充:@METR_Evals 的上限约为 16 小时。 Cog 拥有高达 100 小时的私有企业 eval,并且有信心为此提供财务担保 🤯 METR 数据集:ML 工程、GPU kernel、网络安全 > “METR(2026)使用 GPT-4o 和 GPT-5 的组合,从压缩的 Claude Code 对话记录中估算人类等效时间。这些对话记录来自 7 名 METR 技术人员在 34 个会话上的采集,并标注了人类 ground truth。”rlog 为 0.83 Cog 数据集:真实场景的 java/typescript/python/c# 功能开发、bug 修复、迁移 > “我们通过邀请 Devin 用户回顾近期有代表性的会话,并估算每个已完成会话在没有 Devin 的情况下需要多长时间,从而收集了一个 ground-truth 数据集。我们的数据集包含来自 126 名用户的 258 个会话,覆盖多样化的企业客户。”在留出集上 rlog 为 0.74 这是开创性的真实世界 eval 工作,也是更广泛的前沿代码 eval 发布的第一部分,我非常期待把它写出来。向 @annarmitchell 和 @ryanbai1412 致敬,他们领导了这项不起眼的最后一公里数据收集工作!!
正文
previously on @cognition_labs
Finally! the first eval ship from cog!!!!!!!!!! 👼🏼 To contextualize: @METR_Evals cap out at ~16 hours. Cog has private enterprise evals up to 100hrs, and is confident enough to put a financial guarantee on it 🤯 METR dataset: ML eng, GPU kernels, cybersecurity > "METR (2026) used a combination of GPT-4o and GPT-5 to estimate the human-equivalent times from compressed Claude Code transcripts. These transcripts were collected from 7 METR technical staff on 34 sessions labeled on human ground truth". rlog of 0.83 Cog dataset: real life java/typescript/python/c# feature dev, bugfixes, migrations > "We collected a ground-truth dataset by asking Devin users to review recent representative sessions, and estimate how long each completed session would have taken without Devin. Our dataset consists of 258 sessions from 126 users across a diverse set of enterprise customers." rlog of 0.74 on held out set this is pioneering real world evals work and part 1 of a broader frontier code evals drop that I'm really looking forward to writing up. huge kudos to @annarmitchell and @ryanbai1412 for leading the unglamorous last mile data collection!!在 X 查看被引用的帖子
来源:@swyx · x.com