研究对比人类反应时间与推理模型在160道常识"最佳解释"题上的推理token消耗,发现人类耗时更久的问题模型也想得更久,且双方倾向答错同一批题。多次采样取平均后信号更清晰,GPT-OSS-20B的人机相关性从0.41升至0.55。思维链长度可作为问题难度的粗略信号,但需跨多次运行测量。
Reasoning models struggle on many of the same problems humans do,
The same problems often slow down humans and reasoning models. Reasoning-token count can tell you something about how hard a problem is, but this paper says not to trust the result from just 1 model run.
The researchers compared human response time with how many reasoning tokens models used on 160 commonsense “best explanation” problems.
Questions that took humans longer generally made models think longer too, and humans and models also tended to get the same questions wrong.
The interesting part is that this signal became clearer when models tried different reasoning paths and the results were averaged. For GPT-OSS-20B, the human-model correlation rose from 0.41 to 0.55.
So chain-of-thought length can be a useful rough signal of problem difficulty, but measure it across multiple runs. That makes model effort more useful as an evaluation signal.
来源:@rohanpaul_ai · x.com