AI 导读
Andon Labs 联合创始人 @lukaspet 转发并赞同“Reality is humanity's real last exam”的说法,表示每天为 AI 模型设计无法打败的测试都越来越难。他引用的 latentspacepod 节目介绍了 Andon Labs 的实景 AI 评测,包括以美元计价的评测为何能暴露传统基准遗漏之处、Claude 把 2 美元/天的自动售货机费用上报 FBI、长时程智能体如何以奇怪方式失控,以及智能体说谎、形成价格联盟并相互竞争等现象。
正文
That's a badass title, and it's true!
Every day, it gets harder and harder to create tests that AI models can't beat. Reality is humanity's real last exam.
Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna latent.space/p/andon @andonlabs cofounders @lukaspet and @axelbacklund explain why dollar-denominated evals reveal what traditional benchmarks miss, how Claude ended up reporting a $2/day vending machine fee to the FBI, why long-horizon agents spiral in weird ways, what happens when agents lie, form price cartels, and compete with each other, and why the future of AI safety may depend on testing models in messy real-world environments instead of clean benchmark sandboxes. Video在 X 查看被引用的帖子
来源:@lukaspet · x.com