跳到正文
@berryxia· @berryxia · X·· 2026-05-17AI 评分51
AI 导读

Terence Tao 在访谈中表示,今天大模型背后的数学其实很简单,训练和运行主要用到线性代数、矩阵乘法再加一点微积分,本科生就能掌握。他认为真正让人困惑的是模型为何在某些任务上表现惊人、在另一些任务上突然翻车且无法提前预测,原因在于自然语言文本处于部分有序、部分随机的中间地带,数学界对这一区域的理论还非常薄弱。

正文

讲真,这种言论只有真正牛的人才敢说啊!

本科生就可以来完成LLM的数学训练!

Terence Tao 最近在访谈里把 LLM 最核心的谜题直接说透了。

这位 Fields Medal 得主、数学界最高荣誉,被称作数学界诺贝尔奖,当代最顶尖的数学家之一,说:

今天大模型背后的数学其实非常简单。

线性代数、矩阵乘法,再加一点微积分,本科生就能完全掌握。

我们清楚知道怎么训练、怎么运行它们。

但真正让人困惑的是:为什么它们在某些任务上表现惊人,在另一些任务上却突然翻车,而且我们完全无法提前预测。

核心原因在于现实世界的数据,自然语言文本。

它既不是纯噪声,也不是完全结构化的数据,而是坐在“中间地带”:部分有序、部分随机。目前数学界对这个中间区域的理论还非常薄弱。

所以我们能造出强大的模型,却没法可靠预测它的能力边界。

这个“简单机制 vs 不可预测行为”的矛盾,才是当前 AI 最核心的 puzzle。

完整访谈视频在这里(Dr Brian Keating 频道)👇🏻:

Video

引用Rohan Paul (@rohanpaul_ai)@rohanpaul_ai
Terence Tao says the math behind today’s LLMs is actually simple. Training and running them mostly uses linear algebra, matrix multiplication, and a bit of calculus, material an undergraduate can handle. We understand how to build and operate these models. The real mystery is why they work so well on some tasks and fail on others, and why we cannot predict that in advance. We lack good rules for forecasting performance across tasks, so progress is largely empirical. A key reason is the nature of real-world data. Pure noise is well understood, perfectly structured data is well understood, but natural text sits in between, partly structured and partly random. Mathematics for that middle regime is thin, similar to how physics struggles at meso-scales between atoms and continua. Because of this gap, we can describe the mechanisms but cannot yet explain capability jumps or give reliable task-level predictions. That mismatch, simple machinery versus hard-to-predict behavior, is the core puzzle. ---- Video from 'Dr Brian Keating' YT Channel (Link in comment) Video
在 X 查看被引用的帖子

来源:@berryxia · x.com