NVIDIA 在论文 Staying on Task 中测试 7 个开源模型处理加法、排序等重复任务,发现 128K token 任务的平均准确率比 4K token 任务低 62.8%,最好的模型在最长任务中也只有 17.1% 做到每项全对。模型会看似理解任务却中途丢失位置,尤其在条目没有 ID 时;作者建议给每个条目编号、小批次处理并逐行检查输出。
论文给出长任务可靠性下降的具体数字,并附带可迁移的做法,适合构建长流程智能体时参考。
New Nvidia paper: AI models get sloppier as jobs get longer, even inside their context window, so number every item and split big jobs into small chunks.
Model size didn't guarantee reliability on long, repetitive jobs
Picture an agent updating a huge invoice file line by line. It can read the whole file and still skip a line or update the wrong record.
NVIDIA tested 7 open models on simple, repetitive jobs like adding numbers and sorting lists. Average accuracy was 62.8% lower on 128K-token jobs than on 4K-token jobs.
Even the best model got every item right in only 17.1% of the longest jobs. The models seemed to understand the task but lost their place, especially when items had no ID numbers.
If your agent works through long lists, give every item an ID, process them in small batches, and check every line of output.
– arxiv. org/abs/2609.38712
Title: "Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability"
来源:Rohan Paul · x.com