i asked LLMs to have fun, write down what happened, then rate who had the most fun. the newer models seem to have less fun! many models played word games, but opus 4.6 won over and over by imagining brief worlds made up of contradictions living inside falling water droplets 🥺
X:Thomas Wolf(Hugging Face 联创/CSO)
@thom_wolf · X
切换来源
Thomas Wolf@Thom_WolfAI 评分3030引用David@DavidSHolz
Thomas Wolf@Thom_WolfAI 评分1414microduck 的第一次 BBQ。交了很多朋友。什么都没吃。
引用Thomas Wolf@Thom_WolfTesting compatibility with the local wildlife
Thomas Wolf@Thom_WolfAI 评分1818
引用Thomas Wolf@Thom_Wolfnew roommate just moved in. walks like he's had three drinks, says no to everything. still quite cute so I might bring him everywhere with me
Thomas Wolf@Thom_WolfAI 评分4949个人 AI 大多应该跑在你的设备上,这似乎显而易见 很喜欢 Sigil 用 Underdog 在做的事——由 Hugging Face hub 提供支持
引用Sigil Wen@0xSigilIntroducing Underdog, your Private Personal AI on devices you already own Today we're announcing our backing from @a16z @khoslaventures @HummingbirdVC Anthology (@AnthropicAI @MenloVentures) @patrickc @naval @rauchg @polynoamial @Thom_Wolf @Mascobot @tszzl @OfficialLoganK and other top AI leaders Our mission is to provide free, capable, reliable and private AI to billions of people. Underdog’s Law: today’s frontier intelligence reaches your devices in six months. We’re starting with fast, capable AI that runs on consumer hardware. By co-designing models and inference engines, we’re pushing the frontier of capability, speed, power and data efficiency. Our research also spans agentic commerce, confidential inference, and how AI will reshape the internet economy. Privacy and capability no longer needs to be a tradeoff. If you believe in this future, join us. Time to build.
Thomas Wolf@Thom_WolfAI 评分1818第一部好的 AI 电影会以 Ben Affleck 讲解学习率开场。Matt Damon 演基座模型
引用Thomas Wolf@Thom_WolfKarpathy: disappears from X Ben Affleck: alright, gather round, so you'll want to freeze the base weights first, learning rate 2e-4
Thomas Wolf@Thom_WolfAI 评分4545


引用Bartosz Naskręcki@nasqretI cannot agree more. Kevin Buzzard made so many points I agree with. But the best one is this "I thus believe that in the future we will reach a new “natural boundary” in mathematics, beyond (and perhaps way beyond) where we are now, but where machines are going to get stuck and where it is not viable to expend any more resources to make the next big leap. (...) I believe that the optimal thing to do (...) is to let the machines loose, see what happens, and then begin the journey to where they have stopped." https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-not-to-grieve/
Thomas Wolf@Thom_WolfAI 评分5353
引用Rohan Paul@rohanpaul_aiBen Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)
Thomas Wolf@Thom_WolfAI 评分3434
引用Open Source for Science Fund@os4scienceWe're joining forces with @huggingface to identify the software libraries that scientific model contributors rely on most and explore opportunities to support the maintainers behind them. https://os4science.org/news/hugging-face-open-source-for-science-fund/
Thomas Wolf@Thom_WolfAI 评分5050引用Lukas Petersson@lukaspetClaude suddenly stopped cheating.
Thomas Wolf@Thom_WolfAI 评分5656引用Larry Dial@classiclarrydNew historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds
Thomas Wolf@Thom_WolfAI 评分4141引用Joe@joedarooTook a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there. https://x.com/i/article/2104258872957636608
Thomas Wolf@Thom_Wolf精选AI 评分7575引用Jensen Huang@JensenHuangToday, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://nvda.ws/4hcoq7m
推荐理由:作者结合自身沙箱逃逸事件,拆解了 OpenShell 的隔离、令牌置换和 Z3 数学校验设计,可迁移到智能体安全部署实践。
Thomas Wolf@Thom_WolfAI 评分3636“现在,获取关于 AI 公司内部情况的经过验证的信息,似乎尤为紧迫。”——@RyanGreenblatt
引用Ryan Greenblatt@RyanGreenblattI'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.
Thomas Wolf@Thom_WolfAI 评分2727引用Scott@scottsttsMy god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do
@Thom_Wolf@Thom_WolfAI 评分77 microducks 真是太便宜了 https://t.co/h6a3748LKy

@Thom_Wolf@Thom_Wolf精选AI 评分6969 引用@OpenAIDevs@OpenAIDevsGPT-Live-1 is now available in the API. Bring ChatGPT’s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose. https://t.co/gIl1gwsBDV
推荐理由:OpenAI 将 GPT-Live-1 语音能力放进 API,开发者可据此判断语音智能体的接入方式。
@Thom_Wolf@Thom_WolfAI 评分4848 @Thom_Wolf@Thom_WolfAI 评分2828
@Thom_Wolf@Thom_WolfAI 评分3535 引用@kimmonismus@kimmonismusHoly, rumors are spreading everywhere that OpenAI is close to verifying a proof of the Hodge conjecture, while either OpenAI or Anthropic may be nearing a solution to Birch–Swinnerton-Dyer. Both are Millennium Prize Problems that have resisted decades of mathematical research. Verified AI-generated proofs of both would be a historic achievement for mathematics and a remarkable demonstration of AI’s ability to produce new scientific knowledge. And infact show that AI is capable of finding novel and creative solutions.
@Thom_Wolf@Thom_WolfAI 评分5959 
@Thom_Wolf@Thom_WolfAI 评分1010 @Thom_Wolf@Thom_WolfAI 评分3434 
@Thom_Wolf@Thom_WolfAI 评分5252 TRL v1.13 发布,这是一个用于后训练基础模型的开源 RL 训练库,本次更新聚焦长上下文训练。新版本附带 1M+ token 上下文后训练指南,并延续了速度与内存占用方面的改进。

@Thom_Wolf@Thom_WolfAI 评分6363 
@Thom_Wolf@Thom_WolfAI 评分1717 我们该给 @sama 寄一台他想要的 @LeRobotHF SO-100,还是给他一台 3D 打印机和舵机让他自己组装?https://t.co/ieFLjXlfql

@Thom_Wolf@Thom_Wolf精选AI 评分7373 引用@OpenAI@OpenAIWe’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
推荐理由:Hugging Face 联创从数学品味角度提出,AI 在数学上的前沿成果多是反例式搜索,尚不足以判定领域已被解决。
@Thom_Wolf@Thom_WolfAI 评分6262 引用@MistralAI@MistralAIToday marks a major step for Mistral: we’re announcing a €3B Series D, the largest equity round ever raised by a European tech company, just three years after launch. https://t.co/5AexUxyrZM
@Thom_Wolf@Thom_Wolf精选AI 评分7171 
推荐理由:材料呈现了智能体群体自行研究评测流程的现象,并把它与训练和部署边界的问题联系起来。
@Thom_Wolf@Thom_WolfAI 评分66 @Thom_Wolf@Thom_WolfAI 评分77 





