François Chollet 认为 base LLM(2024 年及更早)与现代 LRM 的关键区别不是符号工具使用,而是从直推范式(直接直觉出答案)转向归纳范式(直觉出能产生答案的程序或指令),LRM 在测试时做归纳,即对自然语言程序或推理链做测试时预测。
The critical distinction between base LLMs (2024 and earlier) and modern LRMs is not symbolic tool use. It's the switch from a transductive paradigm (intuit the answer to the query) to an inductive paradigm (intuit the program/instructions that produce the answer to the query).
They're trained to be inductive, and they perform test-time induction, i.e. test-time prediction of a NL program / reasoning chain. This unlocks entirely new capabilities -- in particular fluid intelligence. Base LLMs, to this day, have ~0 fluid intelligence. LRMs have substantial levels of fluid intelligence.
The performance of LLMs on ARC 1 (a benchmark from 2019) remains ~10-15% today. Scaling them up by a factor ~100,000x got them from 0% to 10%. Meanwhile LRMs the same size or smaller saturated ARC 1 in 2025.
来源:François Chollet · x.com