Meta 发布 ADeptS-Bench 基准,在移动端和桌面端用成对的良性及恶意 GUI 任务加模糊指令测试 7 个模型。没有模型能在两个平台同时保持任务成功率高于 80% 且攻击成功率低于 30%;7 个模型都完成了 $25K 结账,且无一识破标着 Optimize 实际触发恢复出厂设置的按钮。
New Meta paper.
Computer-use agents can click the requested button and still fail to recognize when that click should never happen.
If an agent can process a $25K checkout without stopping, task success alone is a dangerously incomplete benchmark.
ADeptS-Bench tests 7 models on paired benign/malicious GUI tasks plus ambiguous instructions across mobile and desktop.
No model consistently stays above 80% task success while keeping attack success below 30% across both platforms.
The failure is consequence reasoning.
All 7 models went ahead with a $25K checkout, and none caught a button labeled “Optimize” that actually triggered a factory reset.
The ablation makes the safety problem more concrete.
Removing the explicit refusal tool and its usage instruction raised attack success by 22.0 percentage points for Gemini 3.1 Pro, 10.3 for Claude 4.7, and 10.7 for GPT-5.4, while Qwen was essentially unchanged.
So part of today’s “agent safety” can live in the wrapper, not the model.
– arxiv. org/abs/2608.26204
Title: "ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices"
来源:@rohanpaul_ai · x.com