跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 24 天前AI 评分51
AI 导读

一篇论文比较 8 个 LLM 与 18000 多名人类学习者,考察答对难题时是否也掌握其前置的简单技能。人类整体得分 79.6%,其中 72.7% 的正确答案所有受测前置技能也答对;Qwen3-80B-Instruct 得分更高达 92.5%,但该一致比例仅 48.16%。

正文

Humans usually need the foundations before the advanced skill; we need to know the basics before they know the harder thing

But LLMs can get the advanced answer right while missing the foundations underneath it.

The paper compares 8 LLMs with more than 18,000 human learners and asks: if you can solve a harder problem, can you also solve the easier skills underneath it?

Humans were much more consistent.

They scored 79.6% overall, and 72.7% of their correct answers also had every tested prerequisite correct.

QWEN3-80B-INSTRUCT scored higher at 92.5%, but reached that same consistency on only 48.16% of its correct answers.

In a separate test, prerequisite examples were not consistently better than same-skill or similar examples.

Even standard reasoning judges largely missed this pattern.

So high accuracy can hide disconnected pockets of knowledge.

So the paper recommends: evaluate reasoning models with connected sets of easy and hard problems, not isolated benchmark questions alone.

来源:@rohanpaul_ai · x.com