黄仁勋在 Axios 访谈中被问到开源模型公司是否应被允许蒸馏闭源模型,他表示蒸馏、向 AI 和他人学习是智能的基础。他认为 AI 生成的内容正超过人类,未来几年互联网可能 99% 由 AI 生成,AI 系统会持续从其他 AI 系统蒸馏知识。他称 AI 能学习是好事,更聪明的 AI 也可以更安全。
Jensen Huang on "distillation"
In this interview with axios, he was asked this question: "Should open source model companies be allowed to distill closed models"
"Distillation, learning from AI, learning from other people, and learning from other sources of knowledge, is fundamental to intelligence.
We are constantly learning from other people. I am learning from you through the questions you are asking, and you are learning from me. All day long, we are learning from one another. AI also has to learn from something.
The original AI models, whether they were open or closed, were trained on previously created knowledge from the internet. Now, AI is generating more content than humans. In a few more years, the internet could be 99% AI-generated content, and that content will have been created by some form of AI.
As a result, AI systems will constantly be distilling knowledge and intelligence from other AI systems. The fact that AI can learn is a good thing. We want AI systems to be intelligent because a smarter AI can also be a safer AI."
----
From "Axios" YouTube channel, (full video link in comment)
In a new allegation, the U.S. government has accused 6 Chinese AI firms of using large-scale distillation to copy American model capabilities. The NSA, CISA and FBI published joint advisory AA26-251A, naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z .AI as running industrial-scale distillation campaigns against U.S. frontier models since at least late 2024. The agencies say requests moved through a gray market of API proxies called transfer stations, bulk premium subscriptions shared across developer teams, and third-party aggregators that strip account metadata. Moonshot AI allegedly distilled 17 U.S. models, including Anthropic's Claude Fable 5, to train Kimi-K3. They claim DeepSeek's prompts pushed models to write out hidden chain-of-thought steps, which transfers reasoning method rather than finished answers. So they say that DeepSeek's headline training cost only covers the compute it burned. The advisory argues that number is misleading because it leaves out what the training data actually cost, which DeepSeek allegedly obtained by distilling U.S. models rather than producing it through its own research. So the cheap-training story rests on data someone else paid to create. The mitigation section asks U.S. labs to serve suspected distillers subtly degraded responses without telling them. Its detection indicators are behavioral, covering sustained 24/7 usage, new accounts at immediate maximum throughput, and traffic optimized for cache hits. Those patterns also describe an ordinary enterprise agent fleet, leaving each provider to decide which customers receive an undisclosed downgrade.在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com