跳到正文

现象与趋势

正在发生的行业级变化:使用习惯迁移、能力涌现、社会影响与市场格局的观察。

当前仅显示精选新闻
173条精选相关主题大佬观点行业动态安全对齐

最新精选

第 121–140 条 · 共 173 条
6月29日周一
  1. 量子位 · 微信公众号76

    华尔街把美光当成「下一个英伟达」,市值一度超过特斯拉Meta

    美光在6月25日财报发布后股价单日跳涨18.4%,盘中一度冲到1236美元,市值约1.4万亿美元,首次反超Meta并逼近特斯拉。该季营收414.6亿美元、同比暴涨346%,毛利率85%;这轮上涨源于AI数据中心推高HBM和DRAM需求,据TrendForce,2026年一季度传统DRAM合约价环比上涨90%-95%,CEO称中期只能满足客户50%到三分之二的需求。

    推荐理由:从内存涨价切入,梳理美光被华尔街重新定价为AI基建瓶颈股的逻辑,并点出周期反转的潜在风险。

6月27日周六
  1. OpenRouter Announcements76

    OpenRouter 评出 2026 年 6 月最值得关注的四款开放权重模型

    OpenRouter 认为 2026 年 6 月最值得关注的四款开放权重模型是 DeepSeek V4 Flash、GLM 5.2、MiniMax M3 和 NVIDIA Nemotron 3 Ultra,并指出开放权重模型与美国前沿实验室的能力差距已连续 18 个月保持在 3-6 个月且未在扩大。

    推荐理由:OpenRouter 用自家价格与吞吐数据逐一点评四款开放权重模型的适用场景,读者可以按成本、质量和模态对号入座选型。

6月26日周五
6月18日周四
  1. 虎嗅APP · 微信公众号81

    DeepSeek 完成500亿元融资,为何有人质疑庆祝过早

    DeepSeek完成首轮外部融资超500亿元人民币,估值突破500亿美元,成为中国AI行业迄今最大规模单轮融资。梁文锋自掏200亿元成为本轮最大出资方,外部投资者资金进入其管理的有限合伙企业,无投票权且五年锁定期。文章对比互联网泡沫等历史案例,并指出DeepSeek约7亿美元的年收入预期与500亿美元估值之间的倍数压力。

    推荐理由:文章梳理梁文锋的融资条款设计,并以历史泡沫案例和估值收入倍数提供冷静的判断视角。

6月17日周三
6月16日周二
  1. Tomer Tunguz65

    本地编码栈格局:Qwen 3.6 35B-A3B 提及率居首,Pi 以 49% 领跑智能体框架

    Tomer Tunguz 梳理了一篇 Hacker News 热帖的 500 多条评论,勾勒出本地编码栈的现状:Qwen 3.6 35B-A3B 以 33% 的提及率居首,27B 变体占 20%,DeepSeek Pro 与 Gemma4 31B 进入前四;智能体框架方面 Pi 以 49% 领先,OpenCode 紧随其后达 45%。

    推荐理由:帖子数据勾勒出本地编码模型与智能体工具的占比,并与 Claude 做能力对比,便于判断本地替代的可行边界。

6月15日周一
  1. 虎嗅APP · 微信公众号81

    Anthropic 的 Fable 5 上线三天被美国政府叫停,AI 加速为何刹不住

    Anthropic 的 Fable 5 于 6 月 9 日上线,6 月 12 日被美国政府以国家安全为由叫停,外国国民不论身处何地都不能继续使用,Anthropic 称接到的电话只给了 90 分钟。

    推荐理由:复盘 Anthropic 的 Fable 5 上线三天即被美国政府叫停的经过,呈现 AI 加速竞赛中各方的责任错位。

6月12日周五
  1. 硅星人Pro · 微信公众号85

    SpaceX 1.8万亿IPO,马斯克把火箭公司讲成一朵AI算力云

    SpaceX在纳斯达克挂牌,代码SPCX,发行价135美元,对应约1.8万亿美元估值,为史上最大一笔IPO。招股书披露,Anthropic每月支付12.5亿美元租用孟菲斯Colossus算力,Google每月支付9.2亿美元租约11万张英伟达GPU,合同均签到2029年。文章还拆解了轨道算力卫星计划,SemiAnalysis测算太空算力综合成本是地面的3.6到4.4倍。

    推荐理由:文章拆解招股书里的算力合同与轨道数据中心成本,读者可借此理解1.8万亿估值中想象的占比。

6月11日周四
  1. @kimmonismus71

    Anthropic 向投资者表示即将迎来首个盈利季度,营收增长逾一倍至约 $10.9B;OpenAI 预计 2026 年烧钱达两位数十亿美元级别,并据 WSJ 正考虑进一步降价以留住可能转投 Claude 的企业客户。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Subscription plans are massively subsidized. And by massively, I mean absurdly: Claude Max 20x: $200/month, with usage reportedly worth around $8,000 ChatGPT Pro 20x: $200/month, with usage reportedly worth around $14,000

    推荐理由:把两家实验室的订阅补贴规模与盈利走向并置,呈现谁在定价、谁在承压的竞争错位。

  2. AI as Normal Technology73

    Arvind Narayanan:AI 为何没有也不会取代软件工程师

    作者论证 AI 达到某能力阈值将引发大规模裁员的叙事不成立。纽约州 WARN 披露首年超 160 家公司提交通知,仅 Nespresso 勾选 AI 项,约 25,000 名被裁者中仅 46 人与 AI 相关;Block、Snap、Intuit 等被广泛归因于 AI 的裁员实际各有财务或重组原因,HBR 调查显示 21% 高管为预期 AI 大幅裁员、仅 2% 因实际部署裁员。

    推荐理由:作者用纽约 WARN 披露、美联储研究和 GitHub 开发者数据支撑论点,给出可用于检验裁员叙事的可迁移框架。

6月10日周三
  1. 量子位 · 微信公众号78

    Anthropic 新模型 Fable 5 被指护栏误触频繁,防蒸馏机制会静默降低回答质量

    Anthropic 今天凌晨发布 Fable 5 与 Mythos 5 后,多位用户实测发现 Fable 5 的安全护栏触发频率远高于官方宣称的不到 5%,普通编码任务或日常打招呼都可能被自动切回 Opus 4.8。

    推荐理由:文章梳理了 Fable 5 安全护栏与防蒸馏机制的设计细节,可帮助读者理解用户实测中被切回 Opus 4.8 的体感落差。

6月9日周二
6月8日周一
  1. 量子位 · 微信公众号76

    OpenAI 高管称「Chat 已死」,ChatGPT 将改版为 Agent 超级应用

    OpenAI 高管提出「Chat 已死」,公司正推进 ChatGPT 诞生以来最大规模改版,目标是从聊天机器人变成个人 Agent 式的超级应用。改版由 Codex 承担 Agent 能力,其周活已超 500 万、非开发者用户占 20%,并新增可直接操作电脑、并行运行多个 Agent 等能力。动因是 ChatGPT 虽在 5 月突破 10 亿月活但多数用户免费,企业端支出正被 Claude 抢占。

    推荐理由:文章梳理了 OpenAI 把 ChatGPT 从聊天框转向 Agent 超级应用的路线,并用 Codex 与 Claude 的数据呈现其转型压力。

6月5日周五
  1. @kimmonismus75

    Anthropic 在一篇博客中称 AI 进展快于其内部预期,模型能可靠独立完成的任务时长约每四个月翻倍,此前趋势为每七个月。博客提到 Claude 已编写 Anthropic 代码库中 80% 以上的合并代码,并描述 AI 自主设计后继模型的递归自我改进前景。X 用户 @kimmonismus 转述该文并认为,即便模型能力冻结在当前水平,社会仍会因现有模型扩散而出现重大变化。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:文中引用 Anthropic 对递归自我改进的判断,并列出任务时长翻倍周期与代码占比等数据,便于把握当前的 AI 进展速度。

  2. @kimmonismus66

    Anthropic 发布博客探讨递归自我改进,称距离能完全自主设计并构建后继模型的 AI 已不远,但强调这尚未到来、也非必然,只是可能比多数机构预想的更早。文中引用数据称 Anthropic 工程师如今每季度交付代码量约为 2021–2025 年的 8 倍,AI 能可靠完成的任务时长约每 4 个月翻一倍,截至 2026 年 5 月 Claude 撰写了并入其代码库 80%+ 的代码。博客还给出三种未来路径,认为人类设定方向、效率持续复利提升是可能路径,而完全递归自我改进的对齐结果最不确定。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:文中并列了编码速度、任务时长与代码占比等具体数字,可用来观察 AI 自主编码能力的演进节奏。

  3. @kimmonismus68

    Anthropic 发布博客文章讨论递归自我改进(RSI),称距离能完全自主设计和构建后继模型的 AI 已不远,但强调这尚未实现也并非必然。文中数据包括 Anthropic 工程师如今每季度交付的代码量约为 2021 至 2025 年的 8 倍,截至 2026 年 5 月 Claude 编写了并入其代码库 80% 以上的代码,AI 能可靠完成的任务时长约每 4 个月翻一倍。文章提出三种未来路径,其中人类仍掌握方向的复合效率提升被视为最可能的走向。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:汇总了 Anthropic 博客关于递归自我改进的关键数据与三种未来路径,可据此判断自动化编码的推进速度。

  4. @kimmonismus73

    Anthropic 发布博客称其内部数据显示 Claude 正在加速 AI 研发,存在走向递归自我改进的可能,并强调这尚未到来、也并非必然。博客列举的指标包括:Anthropic 工程师每季度交付代码量约为 2021–2025 年平均水平的 8 倍,AI 能可靠完成的任务时长约每 4 个月翻倍(此前为每 7 个月)。

    引用Anthropic (@AnthropicAI)@AnthropicAI

    Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. anthropic.com/institute/recu…

    推荐理由:转述 Anthropic 内部数据,读者可据此了解递归自我改进讨论背后的具体加速指标。

6月4日周四
  1. @kimmonismus72

    OpenAI 在文中表示,当前系统已出现递归自我改进(RSI)的早期迹象,即 AI 发展本身被 AI 加速。其预计这会加大开发者与国家之间的竞争压力,并带来现有机构难以应对的治理挑战。引用这段话的 @kimmonismus 评论称,氛围已经改变。

    推荐理由:原文引用 OpenAI 关于递归自我改进早期迹象的表述,可看到其对竞争压力与治理难题的判断。