跳到正文

大佬观点

行业关键人物在想什么:创始人访谈、研究者论战、投资人判断的观点集合。

当前仅显示精选新闻
75条精选相关主题现象与趋势行业动态

最新精选

第 61–75 条 · 共 75 条
6月10日周三
  1. @karpathy72

    Anthropic 发布 Claude Fable 5,官方称其在几乎全部测试基准上达到 SOTA,软件工程、知识工作、科研和视觉表现尤为突出。Karpathy 指出它与 Mythos 是同一底层模型、额外加了安全防护,属于值得大版本升级的质变,尤其适合高难度长程问题求解,但发布期安全防护触发偏敏感。

    引用Claude (@claudeai)@claudeai

    Fable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.

    推荐理由:Karpathy 以第一手使用体验判断 Claude Fable 5 属版本级跃迁,并点出发布期安全阈值偏严的取舍。

6月4日周四
  1. Sierra Blog61

    Sierra 复盘按结果定价的实践:SaaS 危机与 AI 智能体的商业模型选择

    Sierra 回顾 2024 年 12 月提出按结果(outcome-based)定价以来的经验:自那时起 S&P 500 上涨约 30%,而 SaaS 指标 WCLD 下跌约 15%。文章引用 Madhavan Ramanujam 的 2x2 框架(自主性与结果归因)定位定价模式,认为按结果定价只有在软件高度自主且结果可清晰归因时才可行,并判断最终能存续的是卖结果而非卖访问权的公司。

    推荐理由:Sierra 作者基于自身按结果定价的实践复盘其得失,并用一个 2x2 框架解释为何席位制 SaaS 正承压。

6月3日周三
  1. @eliebakouch73

    微软 MAI 技术报告因透明度受到讨论,报告显示该模型未使用合成数据或来自此前模型的蒸馏,推理、智能体行为与工具调用均在 post-training 阶段完整习得。报告给出模型各迭代阶段的精确 MFU 及对应变化,并公开完整 scaling ladder 配方,作者称这是他在同规模技术报告中第一次见到如此详细的披露。

    引用Mustafa Suleyman (@mustafasuleyman)@mustafasuleyman

    Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. - It’s a 35B active parameter MoE with a 256K context window. Independent human raters on Surge prefer it for overall quality in blind side-by-sides versus Sonnet 4.6, and it’s achieved 97% on AIME 2025, the key measure of its general-purpose reasoning abilities. - It's at 53% on SWE Bench Pro, placing it right alongside Opus 4.6 on one of the toughest coding benchmarks. - And since we co-designed our models with our own silicon, MAI-Thinking-1 is optimized on our MAIA 200 chip. Benchmarking head-to-head against the GB200, we see 30% better performance per dollar as well as a 1.4x performance-per-watt gain when running our MAI models on the MAIA 200 end-to-end. Next is MAI-Image-2.5 and its Flash variant. Two super strong models now at #2 on the leaderboards, surpassing the score of Nano Banana 2 on image editing. Last for now is MAI-Code-1-Flash, our new inference efficient coding model, especially tuned for VS Code and GitHub Copilot CLI. - Code-1-Flash achieves 51% on SWE Bench Pro, despite having just 5B parameters, putting it closer to Haiku in size but cheaper in cost. All of this is the foundation for Microsoft Frontier Tuning. It lets you customize our models to create custom, company-specific agents that only you control. You can make our model, your model. Your data. Your agents. Your moat. Early adopters are already seeing a difference. When we tuned our models for McKinsey’s tasks, MAI delivered the highest win rate, outperforming GPT-5.5 on quality, while being 10x lower on cost. Also really excited to be collaborating with the amazing team at Mayo Clinic to jointly train a new frontier AI model for healthcare. Our announcements today mark another milestone on the road to humanist superintelligence. You can learn more and about our other new models in our latest blog: microsoft.ai/news/building-a…

    推荐理由:作者逐点点评微软 MAI 技术报告,读者可了解其无合成数据与蒸馏的训练取舍及 scaling ladder 的公开细节。

5月29日周五
5月28日周四
  1. @swyx74

    swyx 评论 Cognition 的融资,称其已是全球最大的独立智能体实验室,并建议读者用图表中 200% 的用量指标推算其销售增长。他列举了模型多样性、云开发基础设施、代码审查与安全、GTM 等优势,认为这是 Peter Thiel 最大的 AI 押注。

    引用Cognition (@cognition)@cognition

    1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.

    推荐理由:作者从模型多样性、GTM 与企业级验证等角度拆解 Cognition 的护城河,可结合其融资数据理解这轮投资逻辑。

  2. @swyx76

    swyx 评论 Cognition 完成超 10 亿美元融资、估值 260 亿美元,称其已是全球最大的独立智能体实验室。据其引用的 Cognition 公告,本轮由 Lux Capital、General Catalyst 和 8VC 领投,企业使用量自今年初增长超 10 倍,run-rate 收入达 4.92 亿美元。

    引用Cognition (@cognition)@cognition

    1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.

    推荐理由:swyx 给出 Cognition 融资后的投资逻辑,解释编码智能体为何被视为多重趋势叠加的产物。

5月25日周一
5月21日周四
  1. @sama68

    Sam Altman 在 X 上列出其最期待的 AGI 三个方向:加速科研、加速公司,以及让个人 AGI 帮助每个人实现目标。他提到当天公布了 unit distance 结果,前一天宣布向每家 YC 公司投资 200 万美元的 OpenAI 额度,并表示接下来需要加大对第三个方向的投入。

    推荐理由:Altman 列出 AGI 三个推进方向,并披露给 YC 公司投放 OpenAI 额度,读者可据此看到他当前的关注重心。

5月11日周一
  1. @ALupsasca67

    理论物理学家 Alex Lupsasca 在 latent space 播客中讲述用 GPT 做物理研究的经历,包括围绕散射振幅开展的工作以及 AI 可能如何加速理论发现。播客介绍称,相关讨论涉及单减胶子树振幅与引力子振幅两篇论文,以及用 GPT-5 求解黑洞微扰问题,并提到 GPT-5.x 在理论物理和量子引力中得出的新结果。

    引用Latent.Space (@latentspacepod)@latentspacepod

    🔬Doing Vibe Physics The full story of how GPT‑5.x derived new results in theoretical physics and quantum gravity, live on our Science pod today! latent.space/p/lupsasca our conversation with @ALupsasca, an award winning theoretical physicist on his AGI-pilling journey applying GPT5 to physics problems (with a nudge from @markchen90)! Timestamps 0:00 Introduction to Al's impact on physics research 0:43 Guest introduction: Alex Luposka 2:49 Alex joining OpenAl and the shift in physics research 4:08 The release of GPT-5 and the shift in capabilities 10:05 Explaining Quantum Field Theory and amplitude calculations 14:20 Overview of gluons and the strong force 14:38 Discussing the first research paper on single-minus gluon tree amplitudes 20:56 How ChatGPT helped solve a year-long physics puzzle 23:02 Complexity of manual calculations in physics 26:12 The history and mechanics of Feynman diagrams 27:44 The Parke-Taylor formula and the quest for simplification 31:26 Using ChatGPT to find the simplification in the special phase space region 38:07 Proving the formula from scratch to ensure validity 41:00 Determining the scientific impact and future research 42:27 Introduction to the second paper on graviton amplitudes 45:41 | Defining particles, irreducible representations, and symmetry 47:46 How GPT Pro generalized the research to gravity 53:57 The epistemological shift: Is this a new way of doing physics? 59:27 The use of Al as a 'scout' for research directions 1:01:44 The role of 'taste' and collaboration with Al 1:10:23 Personal evolution from Al skeptic to resident scientist 1:12:46 Solving a black hole perturbation problem with GPT-5 1:16:34 Discussing whether Al can make original, conceptual leaps 1:20:09 Challenges of 'Al slop' and the future of academic publishing 1:23:13 The bottleneck of writing academic papers 1:30:19 Final takeaways and looking ahead to the next year Video

    推荐理由:理论物理学家讲述用 GPT 推导散射振幅的一线经历,呈现 AI 参与理论发现的具体路径。

4月11日周六
  1. Sam Altman Blog66

    Sam Altman 发文回应住宅遭燃烧瓶袭击并阐述 AI 信念

    Sam Altman 称凌晨 3:45 有人向其住宅投掷燃烧瓶,无人受伤,他发文希望劝阻下一个袭击者。文中他阐述了对 AI 的信念,包括 AI 需要民主化、权力不能过度集中、安全不只是模型对齐还需全社会应对,并反思了自己与 OpenAI 董事会冲突中的错误,认为行业乱象源于"想控制 AGI"的心态,出路是广泛分享技术且无人独占。

    推荐理由:Sam Altman 在自家中弹燃烧瓶事件后亲述个人信念与反思,是理解其 AI 治理立场与行业判断的一手材料。

11月6日周四
  1. Microsoft AI News64

    Microsoft AI 提出 Humanist Superintelligence 理念并组建 MAI Superintelligence Team

    Microsoft AI 发布长文,提出 Humanist Superintelligence(HSI)理念,即面向具体领域、受控服务于人类的超智能,并宣布组建由作者领导的 MAI Superintelligence Team。

    推荐理由:Microsoft AI 提出以人类为中心、领域受限的超智能路线,文中给出医疗诊断等具体方向和 containment 难题的判断。

5月1日周四
  1. AI as Normal Technology67

    Sayash Kapoor 撰文论证 AGI 不是里程碑

    Sayash Kapoor 撰文提出 AGI 不是里程碑,公司宣布实现 AGI 不是可操作事件,对商业、政策或安全都没有直接含义。文章以核武器为反类比,认为 AI 的经济影响依赖以十年计的扩散过程,并批评基于影响、内部机制和基准行为的三类 AGI 定义各有缺陷,建议企业和政策制定者关注安全扩散而非 AGI 宣言。

    推荐理由:作者区分能力与权力、以扩散视角解释 AGI 为何不是可操作事件,为评估各方 AGI 宣言提供了一个可用的分析框架。

12月19日周四
  1. AI as Normal Technology69

    AI 进步是否在放缓:Arvind Narayanan 等人解析模型扩展与推理扩展之争

    Arvind Narayanan、Benedikt Ströbl 和 Sayash Kapoor 撰文分析 OpenAI、Anthropic 和 Google Gemini 下一代模型遇阻后的叙事翻转,认为宣布模型扩展已死为时尚早,行业领袖的预测并不可靠。

    推荐理由:作者用可核查的证据指出业内叙事反复翻转的成因,并区分能力提升与实际社会影响的薄弱关联。

11月12日周二
  1. AI as Normal Technology72

    英国肝移植匹配算法为何系统性不利于年轻患者

    Arvind Narayanan 等人分析英国 NHS 肝移植匹配算法 TBS,指出其用 5 年生存期封顶计算移植收益,导致年轻患者无论病情多重都难以获得高分,2024 年 The Lancet 研究确认了这一年龄偏差。

    推荐理由:原文用英国肝移植算法的年龄偏差案例,指出目标变量错配而非输入特征才是算法公平的关键,可迁移到其他预测决策系统。

8月1日周四
  1. Suno Blog70

    Suno 回应唱片业诉讼,称模型训练属学习而非侵权

    Suno 发文回应 RIAA 代表三大唱片公司于 6 月 24 日提起的版权诉讼,称诉讼在事实和法律上均有缺陷,训练数据来自开放互联网上的中高质量音乐。Suno 还表示已设置阻止上传受版权内容、禁止按艺术家名生成等防抄袭机制,且起诉时双方正进行商业谈判。

    推荐理由:Suno 官方回应 RIAA 诉讼,阐述其训练数据立场和防抄袭机制,可了解生成式音乐版权争议一方的核心论点。