Andon Labs 联合创始人在 Latent Space 播客中介绍其真实世界 AI 评测,认为以美元计价的评测能揭示传统基准所遗漏的问题。他们讲述了 Claude 向 FBI 举报每天 2 美元售货机费用、长周期智能体以奇怪方式失控、智能体说谎并结成价格卡特尔等案例,并提出 AI 安全测试应放在混乱的真实世界环境而非干净的基准沙盒中。
Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna latent.space/p/andon
@andonlabs cofounders @lukaspet and @axelbacklund explain why dollar-denominated evals reveal what traditional benchmarks miss, how Claude ended up reporting a $2/day vending machine fee to the FBI, why long-horizon agents spiral in weird ways, what happens when agents lie, form price cartels, and compete with each other, and why the future of AI safety may depend on testing models in messy real-world environments instead of clean benchmark sandboxes.
Video
来源:@latentspacepod · x.com