微软与清华大学的论文发现,给LLM提供结构化的运行视图而非原始对话历史,可将GPT-5.1对智能体故障步骤的精确归因率从3.63%提升至31.35%。长链路智能体运行中早期错误会引发后续症状,直接阅读完整历史容易归错步骤。论文指出智能体调试效果高度依赖运行轨迹的表示方式,而非仅取决于评判模型强弱,目前精确归因率仍仅31.35%,尚未解决。
New Microsoft + Tsinghua Univ paper finds that agent failures are far easier to diagnose when the LLM gets a structured view of the run.
Long agent runs are messy. An early mistake can create later symptoms, so an LLM reading the whole history may blame the wrong step.
Give the LLM a structured run instead of the raw conversation; that raised GPT-5.1 exact localization from 3.63% to 31.35%.
better agent debugging depends heavily on how the run is represented, not just how strong the judge model is.
This is not solved yet; exact localization is still only 31.35%, but make structured traces and explicit failure checks part of the system before asking an LLM to explain what went wrong.
来源:@rohanpaul_ai · x.com