跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-18AI 评分52
AI 导读

字节跳动 Seed 的论文提出 Harness-IF,评测编码智能体在系统提示、工具描述、技能描述、项目文件和用户指令等实际读取位置上的指令遵循能力。在 12 个前沿模型和 60 个多轮编码任务上,当规则与模型默认行为相违背时,所有模型得分都更差。作者认为,若测试规则常与模型默认行为一致,可能高估智能体的可控性。论文链接为 arxiv.org/abs/2608.11727。

正文

New ByteDance paper asks a much better question about AI agents: did the instruction actually change anything?

We might be overestimating how controllable AI agents are because our tests often agree with their defaults.

Want to know whether your agent actually follows instructions? Give it a rule that goes against what it normally does.

Harness-IF tests that harder case: rules that push against a model's default behavior, spread across the places coding agents actually read, including system prompts, tool descriptions, skill descriptions, project files, and user instructions.

Across 12 frontier models and 60 multi-turn coding tasks, every model scored worse on these against-prior rules, i.e. once instructions pushed against its natural defaults.

– arxiv. org/abs/2608.11727

Title: "Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents"

来源:@rohanpaul_ai · x.com