跳到正文

现象与趋势

正在发生的行业级变化:使用习惯迁移、能力涌现、社会影响与市场格局的观察。

当前仅显示精选新闻
201条精选相关主题大佬观点行业动态安全对齐

最新精选

第 181–200 条 · 共 201 条
5月21日周四
  1. @kimmonismus69

    Cursor 发布 Composer 2.5,在 Artificial Analysis 编码智能体指数上得分 62,比上一代 Composer 2 提升 14 分,位列第三,仅次于 Claude Opus 4.7(max)的 66 分和 GPT-5.5(xhigh)的 65 分。标准版每任务成本 0.07 美元、Fast 版 0.44 美元,而上述两款更高分模型分别约为 4.10 和 4.82 美元。该模型仅在 Cursor IDE 和 Cursor CLI 提供,无外部 API,基于 Kimi K2.5 继续训练;推文作者认为性能略好却贵 60 倍已不再划算。

    引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys

    Cursor's new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.7 and GPT-5.5 variants above it. This release puts Composer among the leading coding agent models, something that wasn’t clear for past releases @cursor_ai has released Composer 2.5, the latest model in its Composer line. Composer 2.5 scored 62 on our Coding Agent Index, a 14 point gain over Composer 2 (48). This puts it in third place of our tested agents, behind only Claude Opus 4.7 (max) in Claude Code (66) and GPT-5.5 (xhigh reasoning) in Codex (65). These cost $4.10 and $4.82 per task respectively, ~10x the cost of Composer 2.5 Fast ($0.44) and ~60x the cost of Composer 2.5 standard ($0.07). Key results for Composer 2.5 in Cursor CLI: ➤ Cost-quality Pareto frontier: At $0.07 (standard) and $0.44 (Fast) per task, Composer 2.5 is cheaper than every other agent scoring above 60 on the Index. Medium-effort peers cost $1.24–$2.21 per task; higher-effort variants land 3-4 points above at $4.10–$4.82 ➤ Per-benchmark gains vs Composer 2: +35 points on SWE-Bench-Pro-Hard-AA (12% → 47%), +2 points on Terminal-Bench v2 (64% → 66%), and +3 points on SWE-Atlas-QnA (69% → 72%). At 47%, Composer 2.5's score on SWE-Bench-Pro-Hard-AA is comparable to Claude Opus 4.7 (max) in Claude Code ➤ Among the fastest coding agents: Composer 2.5 Fast runs at an average wall time of 6.7 minutes per task, the third-fastest agent on the Artificial Analysis Coding Agent Index, behind only Claude Opus 4.7 (medium) in Claude Code (5.8m) and GPT-5.5 (medium) in Cursor CLI (6.2m) ➤ Fast mode enables better responsiveness at 6x pricing: Fast runs 30% faster than standard Composer 2.5, but is ~6x the cost per task ($0.44 vs $0.07). Token pricing is 6x higher for Fast: $3.00/$15.00 vs $0.50/$2.50 per million input/output tokens Model details: ➤ Base model: Continued training on @Kimi_Moonshot's open weights Kimi K2.5 as with Composer 2, with Cursor reporting ~85% of total compute from its own additional training and reinforcement learning ➤ Pricing: $0.50/$2.50 per million input/output tokens for the standard variant; $3.00/$15.00 for the Fast variant (the default in Cursor) ➤ Available exclusively in Cursor: both Cursor IDE and Cursor CLI, an externally accessible API is not available Congratulations @cursor_ai and @mntruell on the impressive release!

    推荐理由:推文用每任务成本对比 Composer 2.5 与两个更高分编码智能体,读者可据此权衡编码任务上的性能与花费。

  2. Tomer Tunguz82

    SpaceX 递交 S-1,披露 Starlink、发射与 AI 三大业务数据

    SpaceX 递交 S-1,披露 2025 年 187 亿美元合并营收与 66 亿美元调整后 EBITDA。文件把公司分为 Space、Starlink 和 AI 三个分部,Starlink 贡献 61% 营收、2025 年运营利润 44 亿美元,AI 分部当年投入 64 亿美元建设 COLOSSUS 数据中心并训练 Grok。

    推荐理由:S-1 数据把 SpaceX 拆成卫星、发射与 AI 三块业务,读者可借此比较 AI 算力投入与收入回报的差距。

  3. @AYi_AInotes68

    阿易 AI Notes 引用泄露音频称,扎克伯格在 4 月 30 日全员会上表示 Meta 正用员工的键盘、鼠标、屏幕数据训练 AI,认为员工平均智力高于外包,可更快提升 Llama 的编码能力。作者指出 20 天后 8000 名员工收到裁员邮件,并批评这是把员工当免费高质量训练数据、用完就裁的做法。

    引用More Perfect Union (@MorePerfectUS)@MorePerfectUS

    LEAKED AUDIO: In an all-hands meeting on April 30, Mark Zuckerberg tells employees that he's training AI on them ahead of mass layoffs. "The AI models learn from watching really smart people do things... The average intelligence of the people who are at this company is significantly higher than the average set of people that you can get to do tasks. So if we're trying to teach the models coding, for example, then having people internally build tools or solve tasks that help teach the model how to code, we think is going to dramatically increase our model's coding ability faster than what others in the industry have the capability to do, who don't have thousands and thousands of extremely strong engineers at their company." Video

    推荐理由:引用泄露的全员会音频,呈现 Meta 用员工数据训练 AI 与随后裁员的关联叙事,可借此了解事件背景。

5月20日周三
  1. @berryxia73

    Google 发布 Gemini 3.5 Flash,Artificial Analysis 测试显示其 Intelligence Index 为 55 分,比 Gemini 3 Flash 高 9 分,超过 Grok 4.3 和 Claude Sonnet 4.6,输出速度超 280 tokens/s,比上一代快 70%,幻觉率从 92% 降到 61%。

    引用Berryxia.AI (@berryxia)@berryxia

    兄弟们! 今天已经可以在ZenMux上免费体验Gemini 3.5 Flash 了! 我第一时间用它跑了那个经典的「AI模型递归二叉树生长测试」. 同一个 Prompt ,不同模型画出的树形态完全不一样。(见视频-Prompt见评论区) Gemini 3.5 Flash 从输入提示词到生成完整 HTML 动画网页(树干慢慢长出、分支递归展开、最后随风摇摆),全程只用了 77.56 秒! 整体效果非常惊艳:树形态自然优雅、生长动画丝滑、视频和内容呈现都顶级! 熟悉的老朋友都知道,ZenMux 每次新模型都是 ZeroDelay 首发. Google I/O 2026 今天刚发布,现在立刻就能通过 API 调用! 还有免费额度可以白嫖~ 速度是真的没话说,还完美保留了旗舰级模型的能力。 专为 Agent 设计,在 MCP Atlas、Toolathlon、Finance Agent 等多项榜单直接拿下第一! 多模态理解也极强:MMMU-Pro 83.6%、CharXiv Reasoning 84.2%,全面超越上一代 Gemini 3.1 Pro。 完全兼容主流 API 格式,无需改动现有工具链。 支持按量计费 + Builder 套餐。 👇 直接体验 正式版 → zenmux.ai/google/gemini-3.5-… 免费试用 → zenmux.ai/google/gemini-3.5-… Video

    推荐理由:原文用基准与定价的对比说明 Flash 系列的定位变化,读者可据此重新评估轻量模型的成本预期。

  2. @AYi_AInotes70

    Google 在 I/O 上发布 Gemini 3.5 Flash,称其智能与顶级模型相当但输出速度是其他前沿模型的 4 倍,并当天面向所有人开放。作者认为配套的 Antigravity 平台提供桌面端、CLI 和 SDK 全栈开放,目标是做 Agent 时代的 AWS,而 Spark 个人 Agent 只是示范。

    引用Sundar Pichai (@sundarpichai)@sundarpichai

    Just off stage at #GoogleIO, some highlights from this morning 🧵 Gemini 3.5 Flash is available today for everyone in @antigravity and across our products and APIs. Compared to 3.1 Pro, 3.5 Flash is better across almost all benchmarks with huge progress in coding. It’s also comparable to the best models but very fast (4x faster tokens/ second than other frontier models). And when looking at the intelligence versus output speed, it’s in a league of its own in the top right quadrant.

    推荐理由:把胜负手从模型智力转向智能乘速度乘可部署性,作者给出 Google Agent 基础设施的另一种解读。

5月19日周二
  1. @berryxia68

    Anthropic 宣布收购 SDK 与 MCP server 平台 Stainless,该平台自 Anthropic API 早期起就为其生成几乎全部 SDK。作者认为这不只是技术补全,未来 SDK 形态、MCP 协议走向和开发者必须接受的默认行为都会嵌入 Anthropic 自己的产品哲学与安全策略。他由此担心开发者可用的工具链会越来越窄,只剩一种选择。

    引用Anthropic (@AnthropicAI)@AnthropicAI

    Anthropic is acquiring @stainlessapi, an SDK and MCP server platform that has powered every Anthropic SDK since the earliest days of our API. Read more: anthropic.com/news/anthropic…

    推荐理由:作者把 SDK 与 MCP 工具链的归属变化作为切口,讨论这次收购会如何影响开发者的选择空间。

5月16日周六
  1. @op741865

    作者补发飞书 CLI 的 GitHub 地址(github.com/larksuite/cli)并推荐没装的人试试。被引用的分析提到,该 CLI 于 3 月 28 号开源,一个多月达 10000 Star,期间发布 32 个版本、385 个提交,采用面向日常任务、标准 API 和兜底 API 的三层设计,并配套 Skills 作为 Agent 调用说明书。

    引用歸藏(guizang.ai) (@op7418)@op7418

    飞书 CLI 牛皮啊,发布一个月多点就达到 10000 Star 了! 说明用户和市场相当认可这个动作 最近我们可以发现,越来越多的传统办公产品开始发布 CLI 和 Agent。 AI 时代的 SaaS 软件可能得换个做法了:UI 只是最基本的,接下来还要竞争对 Agent 的适配程度以及覆盖率。在这块,我觉得飞书走得相当靠前。 作为一个 IM 软件,飞书在 AI 时代去做这种开放自己所有能力的 CLI 工具,其实是一种非常不传统互联网的尝试。 这对于之前的互联网产品逻辑和经验来说,是一个非常不应该做的决定。 因为他们这个 CLI 几乎可以控制飞书的所有能力:你可以完全不跟飞书的传统 UI 去交互。只跟 CLI 交互,也可以完成飞书上所有的工作。 传统的 IM 办公软件通常非常复杂,入门门槛相对较高。无论从产品逻辑、UI 设计还是交互设计的角度来看,都没有办法太好地消解这种复杂性。 但是 CLI 工具交付给 Agent 以后,就可以快速消解这种复杂性。用户只需要进行对话,这是非常本能的行为,不需要在繁杂的层级列表 UI 里去寻找功能入口。 我拉了一下数据,他们迭代效率也非常恐怖,它们是 3 月 28 号开源的,一个多月发了 32 个版本、385 个提交。 这说明飞书对这块是非常重视的,投入的人力和精力也非常大。 他们在 CLI 本身的设计上也考虑得非常多,下了很多功夫。主要分为三层: 面向日常任务的快捷命令、开放平台对应的标准 API、兜底的 API 调用。 因为人和 Agent 都不喜欢从 2500 个 API 里去寻找参数,但又需要把这些能力暴露出来,所以他们采用了这种分层的形式。 即使做了分层设计,CLI 本身的内容和 API 依然非常多。所以他们把 CLI 作为工具本身,同时做了很多 Skills 用来充当 CLI 的说明书。 Agent 可以分层、分类型地了解应该如何调用这些 CLI 及其命令。 此外,他们在对 Agent 友好的命令包装上做了很多工作,例如: (a) 内置了 Dry Run (b) 结构化输出 (c) 身份选择、权限检查与风险等级评估 (d) 允许 Agent 在发消息前预览请求 (e) 建立了输出格式的“契约”:将成功或失败的结果、原因以及风险提示都放在结构化数据里。 这样如果出错了,AI 可以非常清楚地进行调试和修改,而不是盲目猜测。 其实现在你如果要创业或者做自己的 Agent,就不需要非得写一个界面。 飞书 CLI 加上 Agent 框架可以完成所有的 Agent 产品常见的操作: 你的聊天界面就是你的 Agent 聊天界面; 你的数据库就是飞书多维表格和文档; 你的用户就是把你拉到组织里的群成员;

    推荐理由:飞书 CLI 开源一个多月获 10000 Star,其分层命令与 Skills 设计为 Agent 调用办公软件能力提供了参考。

5月15日周五
  1. @frxiaobei70

    OpenEvidence 已覆盖 65% 的美国医生,4 月单月覆盖 2700 万次临床场景,平均每位医生每月使用 41 次,基本每个工作日都在用。作者原本以为靠医院系统对接,实际是医生用执业编号自行注册、装在个人手机上,医院并不知情。Mount Sinai 的 AI 负责人将这种现象称为 shadow AI,并称医院后来才追着签企业合作。

    引用OpenEvidence (@EvidenceOpen)@EvidenceOpen

    “We did the hardest thing in the history of American health care. We got the majority of American doctors to all voluntarily adopt a single technology platform.” NBC News on how that happened, what U.S. physicians actually do with OpenEvidence, and how partnerships with NEJM, JAMA, NCCN, and Wiley make it possible.

    推荐理由:原文用覆盖比例与每月使用频次拆解医生自下而上的采用路径,呈现医疗场景中影子 AI 的扩散方式。

5月14日周四
  1. @berryxia67

    Krishna Rao 首次公开做客播客,讲述其加入 Anthropic 两年间公司年化营收从约 2.5 亿美元增至 300 亿美元,并主导募集近 750 亿美元资金。他负责 Anthropic 全部算力的采购、分配与动态调度,已签下超过 1000 亿美元的算力采购承诺。博主由此提出,模型能力趋同之下,算力获取与调配能力或成为决定头部 AI 公司胜负的关键变量。

    引用Patrick OShaughnessy (@patrick_oshag)@patrick_oshag

    Krishna Rao is the CFO of Anthropic, and this is his first podcast appearance. He joined the company two years ago when run-rate revenue was about $250M. Today it is $30B. He has helped raise ~$75B and is responsible for the procurement and allocation of compute. I feel lucky we get to hear what it is like to sit inside a company this consequential at a moment this pivotal. We discuss: - The cone of uncertainty - How he allocates compute across Trainium, TPUs, and GPUs - What investors misunderstand about model companies - Why the returns to frontier intelligence keep rising - Platform vs application and where Anthropic builds its own products - How Anthropic uses Claude internally I have asked my closing question about the kindest thing more than 500 times. Krishna's answer is one I have never heard before. Enjoy! Timestamps: 0:00 Intro 2:38 The Compute Canvas 6:51 The "Cone of Uncertainty" 11:58 Why the Returns to Frontier Intelligence Are So High 16:45 Recursive Self-Improvement 20:20 Scaling Laws 23:30 Sourcing $100 Billion in Compute 28:05 Platform vs. Application Strategy 32:52 Pricing Dynamics 38:48 How Anthropic’s Finance Team Uses Claude 43:24 Raising Capital & Overcoming Investor Skepticism 52:32 Public Perception, Risks, and Government Regulation 57:25 Mythos Release 1:12:33 What Could Derail the AI Revolution? 1:13:47 Biotech and Healthcare 1:15:31 The Kindest Thing Video

    推荐理由:借 Anthropic CFO 首次播客长谈,读者可了解头部 AI 公司的算力采购规模与分配逻辑。

  2. @AYi_AInotes70

    Anthropic 宣布自 6 月 15 日起,付费 Claude 套餐可领取每月专用信用额度,覆盖 Claude Agent SDK、claude -p、Claude Code GitHub Actions 以及基于 Agent SDK 构建的第三方应用。

    引用ClaudeDevs (@ClaudeDevs)@ClaudeDevs

    Starting June 15, paid Claude plans can claim a dedicated monthly credit for programmatic usage. The credit covers usage of: - Claude Agent SDK - claude -p - Claude Code GitHub Actions - Third-party apps built on the Agent SDK

    推荐理由:把订阅额度改按 API 计费的变化与重度用户的账单涨幅放在一起,便于读者重估 agent 自动化的成本。

  3. @AYi_AInotes70

    Google 威胁情报组(GTIG)确认检测到首个已知的、由 AI 独立开发并被实际部署的零日漏洞野外利用案例,称攻击者原计划发动大规模攻击,其主动反制可能阻止了这一行动,该发现收录在其关于 AI 驱动威胁的新报告中。

    引用News from Google (@NewsFromGoogle)@NewsFromGoogle

    The Google Threat Intelligence Group has detected the first known instance of a threat actor using an AI-developed zero-day exploit in the wild. While the attackers planned a wide-scale strike, our proactive counter-discovery may have prevented that from happening. This finding is part of our new report on AI-powered threats.

    推荐理由:Google 披露 AI 开发的零日漏洞已出现在野外攻击中,读者可据此看到攻击生成门槛与现有检测手段之间的落差。

5月12日周二
  1. @AYi_AInotes76

    Mini Shai-Hulud 供应链攻击已从 TanStack 扩散至 UiPath、Mistral AI 相关包,共 205 个制品被毒化,攻击者在 6 分钟内发布 84 个恶意版本,Socket 在 6 分钟内全部标记。

    引用Theo - t3.gg (@theo)@theo

    I hope you guys understand that this is going to keep getting worse

    推荐理由:梳理了这轮 npm 供应链攻击的传播链条和被绕过的签名验证,读者可据此理解 AI 开发流面临的风险。

5月8日周五
  1. Simon Willison80

    Mozilla 用 Claude Mythos 预览版加固 Firefox,月修复漏洞从 20-30 个增至 423 个

    Mozilla 借助 Claude Mythos 预览版定位并修复了 Firefox 中的数百个漏洞。其安全漏洞修复量从 2025 年每月约 20-30 个跳升至 4 月的 423 个,其中包含一个存在 20 年的 XSLT bug 和 <legend> 元素里一个 15 年的 bug。文中还提到,harness 的许多尝试被 Firefox 现有的纵深防御措施拦截。

    推荐理由:文中对比了 AI 生成安全漏洞报告从噪声到可用的转变,并给出 Mozilla 月修复漏洞数量跳升的具体数字。

5月6日周三
4月30日周四
4月29日周三
  1. AI前线 · 微信公众号78

    GitHub Copilot 转向按量计费、Claude Code 限制 Opus,AI 写代码成本逼近程序员工资

    GitHub 宣布自 2026 年 6 月 1 日起 Copilot 从按请求计费转为基于 GitHub AI Credits 的按使用量计费,Opus 4.7 倍率将从 7.5 倍升至 27 倍;Anthropic 同期限制 Claude Code 中 Opus 模型,Pro 用户需额外付费。

    推荐理由:原文梳理 Copilot 与 Claude Code 转向按量计费的细节,并引入 token 成本与程序员工资的比较框架。

4月6日周一
  1. Linear Now63

    Linear 公布 2026 年 3 月 24 日安全事件复盘:权限过滤漏洞致私有团队数据越权可见约一小时

    Linear 官方发布事故复盘,3 月 24 日 12:07 至 1:10(UTC)部署的代码变更导致同一 workspace 内私有团队数据可能对包括 guest 在内的其他成员可见。

    推荐理由:官方复盘给出完整时间线、变量遮蔽根因和权限测试缺口,可作权限层事故响应的参考样本。

5月1日周四
  1. AI as Normal Technology67

    Sayash Kapoor 撰文论证 AGI 不是里程碑

    Sayash Kapoor 撰文提出 AGI 不是里程碑,公司宣布实现 AGI 不是可操作事件,对商业、政策或安全都没有直接含义。文章以核武器为反类比,认为 AI 的经济影响依赖以十年计的扩散过程,并批评基于影响、内部机制和基准行为的三类 AGI 定义各有缺陷,建议企业和政策制定者关注安全扩散而非 AGI 宣言。

    推荐理由:作者区分能力与权力、以扩散视角解释 AGI 为何不是可操作事件,为评估各方 AGI 宣言提供了一个可用的分析框架。

4月15日周二
  1. AI as Normal Technology65

    Arvind Narayanan 发布长文论文《AI as Normal Technology》,将 AI 视为常态技术

    Arvind Narayanan 发布超过 1.5 万字的新论文,提出将 AI 视为“常态技术”的世界观,认为其影响类似电力和互联网等通用技术,变革性经济与社会影响将以数十年为时间尺度缓慢展开。

    推荐理由:作者提出 AI 是像电力一样的常态技术,经济影响将以数十年计缓慢扩散,并给出风险与政策的替代框架。