


关注 AI 研究者、开发者与机构的动态



“现在,获取关于 AI 公司内部情况的经过验证的信息,似乎尤为紧迫。”——@RyanGreenblatt
I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.
The economics of a Neolab. A neolab is loosely defined as a startup of AI researchers who raises a lot of money pre-production to be able to finance GPU compute to take on a large AI problem. To buy 1000 GB300s or ~14 NVL72 racks will set you back $125-150M for 3yrs with 15-30% upfront. That’s about ~2-2.5MW. Thats about enough to do 10^25 flops a quarter and get to a GPT-4 level model which is 1-2 OOMs off frontier for pretraining. If you post-train on a great open source model, you have a better chance of getting to frontier. The risks are a) you need to spend millions on RL environments too and b) being lapped by another model release while being tied to a base model. For this to payback, you need to give your customers a better and ideally cheaper inference service than a base model and serve them for long enough to recoup your large investment. Even at 50% margin on inference, to recoup $10M in training means serving ~10T tokens (!) if you price like Fable / Astra given a standard cache read / input / output split ($2/M blended). And you have to justify being better than a release like Opus 5.5 which is even cheaper. Often, you end up charging your customers a huge premium in terms of platform fees and compute fees on top of pure inference. Meanwhile, every hour you’re not utilizing your GPUs you are burning money so you typically resell this compute back to a broker or run inference for open models / resell spot instances. At below a ~60% utilization on spot, you will still lose money. Add to that insane cost of talent. So what can you do with the compute? - Not play the model game at all. - Play an entirely different model game (Jev, World Labs) that if big labs played, would either a) cannibalize their business or b) be incrementally not significant revenue c) would cause too much distraction from the main main thing - Acquire a proprietary data set (Peridodic Labs) in enough volume in a domain of usefulness to eclipse frontier quality. Often happens in robotics, biology, chemistry. If you do overcome the challenge of building a model that is useful and well priced beyond big labs models, given the huge price of compute, you still need to play in an area where the revenue / compute ratio is signficant and market demand is large enough to payback your compute spend. It is a difficult game.
AI 产业的目的应该是生产工具,在人类手中改善人类的繁荣与福祉。 而不应是创造一个人类种族的"后继物种"。抱有这种想法,实际上就是当下及未来所有人类的敌人。
Smooth Operator
CI has become the top bottleneck of every engineering team I talk to (including Lindy). Our CI spend has become stratospheric.
MiniMax's latest text model, M3.1-Flash-Preview, debuts today on MiniMax Code. Built for everyday development, it's fast, reliable, and ready for real work, from quick bug fixes to full features.
Parse 5 为常见格式的企业文档提供最优价格。它可解析并返回: ✅ 表格 ✅ 图像 ✅ 边界框 ✅ 流程图 ✅ 以及更多 看看它的表现 👇
如果乔布斯来做 Muse 及硬件 会如何做 https://x.com/i/article/2104186692840427521
我最喜欢的 Paul Graham 文章《如何做出伟大的工作》,浓缩在不到 200 秒的视频里 (感谢 opus)
从 DeepSeek 官方 API 处统计的数据来看,约有 60% 的 DeepSeek Harness 用户使用了至少一个第三方插件。第三方插件是 DeepSeek Harness 用户体验中最具特色且不可缺少的一部分。DeepSeek Harness 团队将持续支持第三方插件生态的繁荣发展,并推动插件 API 趋于稳定,在将来减少和尽量避免破坏性更新。 接下来的几天我个人将每天推荐一个优质的 DSH 第三方插件,欢迎 DSH 插件作者在本 thread 下自荐。我会结合插件质量及后台实际统计到的插件使用量择优推荐。 DeepSeek Harness 团队祝大家中秋快乐阖家幸福! (注:在用户使用官方 API 及模型时,DSH 会向官方 API 上报实际使用的插件包名和版本。此类上报不额外消耗 tokens。)
HuggingFace 上开源了一个用 Claude Opus 5.5 生成的模拟医患对话数据集,覆盖 2194 种疾病,平均每段 33 轮,从描述症状、检查聊到治疗和后续安排。
Songs are my superpersuasion failure mode. They can elicit emotions in me like little else... This might be how I flip. Once AI creates pieces more beautiful than Debussy's I'll have no other option than to kneel before the Divine...
周末动画总是最棒的。📺✨ (引用推文:@awesome_visuals:厨师的复仇。由 MiniMax H3 Max @Hailuo_AI 制作)
the chef's revenge. made with MiniMax H3 Max @Hailuo_AI
claude just generated this 2 minute video about the history of Google and it goes so goddamn hard
claude 刚刚生成了这段 2 分钟的关于 Google 历史的视频,简直太炸了
ProgramDistill 从交互式 Web 应用,到可验证、参考引导的 SWE 任务 论文:https://huggingface.co/papers/2609.18805
My god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do
开源 RL 环境就是赢。当然是在 Hugging Face 上!
was looking for a quiet weekend but xiaomi dropped their rl envs repo last night to put in perspective, if you have to buy some tasks like this its usually hundred to thousand dollars per task so this repo is literally worth millions https://huggingface.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
270M+ miles. 841 fewer injury-causing crashes. Our latest safety data shows the Waymo Driver continues to make roads safer for everyone. Compared to human drivers across 5 territories, it reduced: 📉 Injury crashes by 82% 📉 Serious injury crashes by 95% Full data: https://waymo.com/safety/impact
当你拥有 astra、一堆未读的 Slack 消息、无限循环播放的《罗马松树》第 21 分钟,还有一杯深夜的健怡可乐,还需要什么呢
It’s now easier to build plugins for Claude. We built a new portal to submit your plugin, track review, and see usage. Plugins package MCP and skills, and are becoming the way to build for Claude. MCP usage across Claude products is up 110x this year! https://claude.com/blog/build-plugins-for-claude
两台 Mac mini 等你来拿!!! 现在报名,9 月 28 日前提交参赛作品: https://luma.com/zhkhsnpa
Krea AI 宣布 Krea Agent 支持 motion references 功能,可从任意视频片段中提取动作,并赋予新的角色、场景或元素。功能现已开放使用。
To our most astute listeners, if you've noticed a fresh change in our sound— you're spot on. Our AI hosts would love to explain what's happening (and what you can look forward to)!
致我们最敏锐的听众,如果你注意到我们的声音有了新变化——你说对了。 我们的 AI 主播很乐意解释正在发生什么(以及你可以期待什么)!
Claude Code 现在会在你任务进行到一半触及 5 小时限额时,尝试找到一个优雅的停止点,而不是在编辑中途被切断。它会从你的每周限额中提取一小笔固定额度,用来尽量收尾手头的工作。