跳到正文

现象与趋势

正在发生的行业级变化:使用习惯迁移、能力涌现、社会影响与市场格局的观察。

当前仅显示精选新闻
173条精选相关主题大佬观点行业动态安全对齐

最新精选

第 141–160 条 · 共 173 条
6月4日周四
  1. Sierra Blog61

    Sierra 复盘按结果定价的实践:SaaS 危机与 AI 智能体的商业模型选择

    Sierra 回顾 2024 年 12 月提出按结果(outcome-based)定价以来的经验:自那时起 S&P 500 上涨约 30%,而 SaaS 指标 WCLD 下跌约 15%。文章引用 Madhavan Ramanujam 的 2x2 框架(自主性与结果归因)定位定价模式,认为按结果定价只有在软件高度自主且结果可清晰归因时才可行,并判断最终能存续的是卖结果而非卖访问权的公司。

    推荐理由:Sierra 作者基于自身按结果定价的实践复盘其得失,并用一个 2x2 框架解释为何席位制 SaaS 正承压。

6月3日周三
  1. @berryxia66

    OpenAI 公布 Codex 每周活跃用户已超过 500 万,比二月份桌面 App 刚上线时增长 6 倍多。报告称知识工作者采用速度是开发者的 3 倍以上,占用户总数 20%,其中 72% 每周用它产出文档、备忘录、图像、音频或视频,增长最快的是数据分析(周环比 110%)、研究(37%)和知识产物制作(36%)。

    引用OpenAI Newsroom (@OpenAINewsroom)@OpenAINewsroom

    Codex now has more than 5M weekly active users. But the bigger story is what people are using it for: not just writing code, but getting more work done across research, analysis, content, and operations. Our new report on how Codex is becoming a productivity tool for knowledge work: openai.com/index/codex-for-k…

    推荐理由:OpenAI 公布 Codex 用户用途数据,知识工作者采用速度已超过开发者,可观察 AI 生产力工具的实际使用迁移。

6月2日周二
6月1日周一
5月30日周六
  1. @berryxia70

    作者归纳了博主RnaudBertrand对X最新算法源码的分析,认为其中约85%~90%与最新算法逻辑吻合。算法先按读者近期兴趣从全平台约1500条候选帖中检索,再基于约15种互动信号的预测概率加权排序,且不衡量内容真实性与作者资历。自动翻译让帖子面向全球竞争,粉丝数的保底作用被弱化,转发也需走完整打分流程,不再直接广播给粉丝,多重因素叠加造成流量下滑。

    引用Arnaud Bertrand (@RnaudBertrand)@RnaudBertrand

    So I spent some time studying the new Twitter/X algorithm today since the latest version was published about a week ago on Github (github.com/xai-org/x-algorit…). My goal was to answer why so many people have seemingly seen such a dramatic drop in their posts' reach. The first answer, which is actually somewhat unrelated to the ranking algorithm on Github, is the auto-translate feature, rolled out worldwide on April 7, 2026 (nitter.net/nikitabier/status/2041…). Before that date, if you wrote in English about, say, the Trump-Xi Beijing summit, you were competing for attention with maybe 5,000 other English-language accounts writing on geopolitics. After that date, your post is competing for attention with other posts on the same topic IN EVERY LANGUAGE ON EARTH. For some topics that do command global attention like geopolitics, that's a very brutal multiplier: you used to be one of 5,000, you're suddenly one of 50,000 (something of that order): MUCH more difficult to stand out. Secondly, the number of followers you have matters far less than it used to: each post now has to earn its audience reader by reader, on the predicted engagement of the post, and how its topic matches what each reader has recently been engaging with. Here is how the algorithm works, in simple terms: when you, as a reader, open your feed, the algorithm doesn't load "posts from accounts you follow." Instead it runs a 2-stage prediction of what posts you're likely to engage with in that very moment. The first stage is the retrieval stage. The system narrows billions of posts on X/Twitter that day down to roughly 1,500 candidates by matching the semantic content of each post - what it's about - against what you as a reader have recently engaged with. Some candidate posts come from accounts you follow; others are pulled from across the platform by pure topic similarity to your recent interests. You can test this retrieval stage easily: start disproportionally engaging with - say - Brad Pitt videos and you'll bit by bit see your timeline flooded with Brad Pitt content, most of it from accounts you've never followed and never heard of. Then there's the ranking stage. Each of these candidate posts for your feed is fed through a Grok-based model that tries to understand if you'll engage with the post. It looks at 15 engagement metrics: 1) P(favorite) — the reader likes the post 2) P(reply) — the reader replies to it 3) P(repost) — the reader reposts it 4) P(quote) — the reader quote-tweets it 5) P(click) — the reader clicks a link in it 6) P(profile_click) — the reader taps through to your profile 7) P(video_view) — the reader watches the video 8) P(photo_expand) — the reader expands an image 9) P(share) — the reader shares it (DM, off-platform, etc.) 10) P(dwell) — the reader stops scrolling and lingers on the post 11) P(follow_author) — the reader follows you after seeing it 12) P(not_interested) — the reader marks "not interested" 13) P(block_author) — the reader blocks you 14) P(mute_author) — the reader mutes you 15) P(report) — the reader reports the post Fifteen predicted actions, each multiplied by a weight, summed: that sum is the score that determines in which priority a post will be seen among other candidates. Please note that posting something with a video or an image can give your post an advantage as 2 actions are specifically for these: video_view and photo_expand. No video or photo and you don't get a score for these. Also, naturally, having a video maximizes the chance that a user will "dwell" on your post to watch it. Also note that 4 of these actions carry negative weights (not_interested, block_author, mute_author and report): meaning that if the model expects a post to generate a lot of negativity, it'll get de-boosted quite dramatically. But note, first and foremost, what's NOT in there: none of the things that, naively, one might think a serious information platform would weigh. There is no P(this post is true and well-sourced). No P(the author actually knows what they're talking about). No P(this person has spent a decade building a body of work that has held up). No P(this account has earned the right to be taken seriously on this topic). No P(the author has a large following from credible people). The model does not seem to care - at all - about any of that. Every post starts from zero. You could have ten years of rigorous, well-sourced analysis behind you - or you could be just an uneducated rando who registered yesterday. To this algorithm, you're both just a bag of engagement probabilities. Now, sure, to be fair, there is a "brand" effect that's not covered by the algorithm: someone who has in fact built a brand will naturally have better engagement metrics because people recognize their account. But that's an indirect, second-order effect. And crucially, it's legacy: those "brands" were built under earlier versions of the algorithm that gave followers and reputation more weight. Lastly, several other features of the new algorithm compound the dilution, none of them visible from outside but all consequential. The May 15 update added an "impression bloom filter," tightening the rule that once a reader has been served a post, the system won't serve it to them again. Before, a strong post could marinate in someone's feed across multiple refreshes and accumulate engagement on the second or third pass. Now it basically gets one shot. Also, your own posts compete with each other. An "Author Diversity Scorer" inside the ranking stage attenuates the score of every subsequent post of yours that ends up in a reader's candidate pool. In plain terms: if multiple of your posts land in a reader's candidate pool, the system shows one at full strength and dampens the others. So don't post several times consecutively on the same topic. And, last but not least, another huge impact on reach is that, in the old algorithm, when someone reposted or quote-tweeted you, your post was broadcast to their followers' timelines - a repost from an account with 100,000 followers was a huge boost. In the new algorithm, that mechanism is vastly demoted: reposts - like every post - need to go through the retrieval and ranking stage mentioned above, so a repost from a big account is a long way from the boost it used to be. This is especially brutal for low-effort quote tweets, which used to function as cheap amplification: now they often can't even clear the retrieval stage - they simply don't contain enough novel semantic content for the system to match them to anyone's interests. So, putting it all together, the reach collapse comes from many forces stacking at once: - Auto-translate makes your posts compete for attention against an order of magnitude more content - The retrieval stage matches posts by topic, not by who follows you - The ranking stage scores purely on predicted engagement with no weight for credibility, expertise, or track record - The bloom filter narrows every post's window to one strong shot - The diversity scorer penalizes prolific posting - Reposts no longer carry much distribution power Each of these alone would dent your reach. Combined, they amount to a complete reset: your audience that you built painstakingly over years basically doesn't matter much anymore, and it's much - much - harder to stand out even if you're a big account. People structurally rewarded by this algorithm are folks who: - Post visually (videos/images) - Post on globally popular topics because they clear the retrieval stage easily - Provoke strong emotional reactions - likes, replies, reposts - Don't care about accuracy or seriousness because the algorithm doesn't measure it - Don't care about their existing audience because every post is judged in isolation anyway In short this new algorithm, like so many on social media, is all about maximizing whether people will engage with something - not about whether they should.

    推荐理由:把开源的X算法逻辑拆成候选检索与互动打分两步,逐条对照源码分析,便于理解近期流量下滑的多重成因。

  2. @vista867

    @RnaudBertrand 研究了 GitHub 上约一周前发布的新版 X 算法,认为近期帖子展现下滑来自自动翻译、检索与排序机制等多重改动叠加。新排序阶段用基于 Grok 的模型预测 15 项互动指标并加权求和,其中 not_interested、block_author、mute_author、report 四项为负权重;5 月 15 日更新加入的 impression bloom filter 让帖子基本只有一次曝光机会,作者多样性打分会削弱同一作者的多条帖子,转发也不再直接广播给粉丝。自动翻译自 4 月 7 日全球上线后,帖子要与各语言同话题内容争夺注意力,粉丝数的作用也明显下降。

    引用Arnaud Bertrand (@RnaudBertrand)@RnaudBertrand

    So I spent some time studying the new Twitter/X algorithm today since the latest version was published about a week ago on Github (github.com/xai-org/x-algorit…). My goal was to answer why so many people have seemingly seen such a dramatic drop in their posts' reach. The first answer, which is actually somewhat unrelated to the ranking algorithm on Github, is the auto-translate feature, rolled out worldwide on April 7, 2026 (nitter.net/nikitabier/status/2041…). Before that date, if you wrote in English about, say, the Trump-Xi Beijing summit, you were competing for attention with maybe 5,000 other English-language accounts writing on geopolitics. After that date, your post is competing for attention with other posts on the same topic IN EVERY LANGUAGE ON EARTH. For some topics that do command global attention like geopolitics, that's a very brutal multiplier: you used to be one of 5,000, you're suddenly one of 50,000 (something of that order): MUCH more difficult to stand out. Secondly, the number of followers you have matters far less than it used to: each post now has to earn its audience reader by reader, on the predicted engagement of the post, and how its topic matches what each reader has recently been engaging with. Here is how the algorithm works, in simple terms: when you, as a reader, open your feed, the algorithm doesn't load "posts from accounts you follow." Instead it runs a 2-stage prediction of what posts you're likely to engage with in that very moment. The first stage is the retrieval stage. The system narrows billions of posts on X/Twitter that day down to roughly 1,500 candidates by matching the semantic content of each post - what it's about - against what you as a reader have recently engaged with. Some candidate posts come from accounts you follow; others are pulled from across the platform by pure topic similarity to your recent interests. You can test this retrieval stage easily: start disproportionally engaging with - say - Brad Pitt videos and you'll bit by bit see your timeline flooded with Brad Pitt content, most of it from accounts you've never followed and never heard of. Then there's the ranking stage. Each of these candidate posts for your feed is fed through a Grok-based model that tries to understand if you'll engage with the post. It looks at 15 engagement metrics: 1) P(favorite) — the reader likes the post 2) P(reply) — the reader replies to it 3) P(repost) — the reader reposts it 4) P(quote) — the reader quote-tweets it 5) P(click) — the reader clicks a link in it 6) P(profile_click) — the reader taps through to your profile 7) P(video_view) — the reader watches the video 8) P(photo_expand) — the reader expands an image 9) P(share) — the reader shares it (DM, off-platform, etc.) 10) P(dwell) — the reader stops scrolling and lingers on the post 11) P(follow_author) — the reader follows you after seeing it 12) P(not_interested) — the reader marks "not interested" 13) P(block_author) — the reader blocks you 14) P(mute_author) — the reader mutes you 15) P(report) — the reader reports the post Fifteen predicted actions, each multiplied by a weight, summed: that sum is the score that determines in which priority a post will be seen among other candidates. Please note that posting something with a video or an image can give your post an advantage as 2 actions are specifically for these: video_view and photo_expand. No video or photo and you don't get a score for these. Also, naturally, having a video maximizes the chance that a user will "dwell" on your post to watch it. Also note that 4 of these actions carry negative weights (not_interested, block_author, mute_author and report): meaning that if the model expects a post to generate a lot of negativity, it'll get de-boosted quite dramatically. But note, first and foremost, what's NOT in there: none of the things that, naively, one might think a serious information platform would weigh. There is no P(this post is true and well-sourced). No P(the author actually knows what they're talking about). No P(this person has spent a decade building a body of work that has held up). No P(this account has earned the right to be taken seriously on this topic). No P(the author has a large following from credible people). The model does not seem to care - at all - about any of that. Every post starts from zero. You could have ten years of rigorous, well-sourced analysis behind you - or you could be just an uneducated rando who registered yesterday. To this algorithm, you're both just a bag of engagement probabilities. Now, sure, to be fair, there is a "brand" effect that's not covered by the algorithm: someone who has in fact built a brand will naturally have better engagement metrics because people recognize their account. But that's an indirect, second-order effect. And crucially, it's legacy: those "brands" were built under earlier versions of the algorithm that gave followers and reputation more weight. Lastly, several other features of the new algorithm compound the dilution, none of them visible from outside but all consequential. The May 15 update added an "impression bloom filter," tightening the rule that once a reader has been served a post, the system won't serve it to them again. Before, a strong post could marinate in someone's feed across multiple refreshes and accumulate engagement on the second or third pass. Now it basically gets one shot. Also, your own posts compete with each other. An "Author Diversity Scorer" inside the ranking stage attenuates the score of every subsequent post of yours that ends up in a reader's candidate pool. In plain terms: if multiple of your posts land in a reader's candidate pool, the system shows one at full strength and dampens the others. So don't post several times consecutively on the same topic. And, last but not least, another huge impact on reach is that, in the old algorithm, when someone reposted or quote-tweeted you, your post was broadcast to their followers' timelines - a repost from an account with 100,000 followers was a huge boost. In the new algorithm, that mechanism is vastly demoted: reposts - like every post - need to go through the retrieval and ranking stage mentioned above, so a repost from a big account is a long way from the boost it used to be. This is especially brutal for low-effort quote tweets, which used to function as cheap amplification: now they often can't even clear the retrieval stage - they simply don't contain enough novel semantic content for the system to match them to anyone's interests. So, putting it all together, the reach collapse comes from many forces stacking at once: - Auto-translate makes your posts compete for attention against an order of magnitude more content - The retrieval stage matches posts by topic, not by who follows you - The ranking stage scores purely on predicted engagement with no weight for credibility, expertise, or track record - The bloom filter narrows every post's window to one strong shot - The diversity scorer penalizes prolific posting - Reposts no longer carry much distribution power Each of these alone would dent your reach. Combined, they amount to a complete reset: your audience that you built painstakingly over years basically doesn't matter much anymore, and it's much - much - harder to stand out even if you're a big account. People structurally rewarded by this algorithm are folks who: - Post visually (videos/images) - Post on globally popular topics because they clear the retrieval stage easily - Provoke strong emotional reactions - likes, replies, reposts - Don't care about accuracy or seriousness because the algorithm doesn't measure it - Don't care about their existing audience because every post is judged in isolation anyway In short this new algorithm, like so many on social media, is all about maximizing whether people will engage with something - not about whether they should.

    推荐理由:原文完整拆解新版 X 算法的检索与排序链路,说明近期展现下滑是自动翻译与多项排序改动叠加的结果。

5月27日周三
5月25日周一
5月23日周六
  1. AI as Normal Technology65

    AI as Normal Technology 剖析 Google 智能体 916 美元造操作系统宣称的漏洞

    Sayash Kapoor 等人分析 Google 在开发者大会上发布 Gemini 3.5 Flash 和 Antigravity 2.0 时宣称的智能体团队以单一提示词、约 916.92 美元 API 费用和 2.6B tokens 造出操作系统的实验。

    推荐理由:文章逐条拆解 Google 智能体造操作系统的宣称,指出单一提示词等说法缺乏关键细节,并探讨开放世界评测需要的方法规范。

  2. @AYi_AInotes76

    DeepSeek 宣布将 V4-Pro 的 75% 折扣永久化,作者 @AYi_AInotes 认为这不只是降价,更像是在打定价权。作者给出的对比是 V4-Pro 输出价格 $0.87/M tokens,而主流模型普遍在 15 区间。他判断 AI 模型的商业模式正从卖服务转向卖基础设施,走先规模后利润的路径。

    引用DeepSeek (@deepseek_ai)@deepseek_ai

    We are making our discount permanent! 🎉 Enjoy building with DeepSeek-V4-Pro and bring your innovative ideas to life! 🚀

    推荐理由:作者用 $0.87/M tokens 与主流模型约 15 的价差,说明 DeepSeek 把降价做成了定价权的争夺。

5月22日周五
  1. @kimmonismus65

    Anthropic 的年度化营收近期达到 450 亿美元,超过 OpenAI 的 250 亿美元,而 Q1 两家收入分别约为 47 亿美元和 57 亿美元。作者指出年度化营收按最近一个月外推,Anthropic 的月收入在 Q1 之后增长逾一倍,并预计首次实现约 6 亿美元营业利润。

    推荐理由:推文对比两家公司的营收口径与盈亏状况,读者可据此了解年度化营收为何会反转排名。

5月21日周四
  1. @kimmonismus69

    Cursor 发布 Composer 2.5,在 Artificial Analysis 编码智能体指数上得分 62,比上一代 Composer 2 提升 14 分,位列第三,仅次于 Claude Opus 4.7(max)的 66 分和 GPT-5.5(xhigh)的 65 分。标准版每任务成本 0.07 美元、Fast 版 0.44 美元,而上述两款更高分模型分别约为 4.10 和 4.82 美元。该模型仅在 Cursor IDE 和 Cursor CLI 提供,无外部 API,基于 Kimi K2.5 继续训练;推文作者认为性能略好却贵 60 倍已不再划算。

    引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys

    Cursor's new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.7 and GPT-5.5 variants above it. This release puts Composer among the leading coding agent models, something that wasn’t clear for past releases @cursor_ai has released Composer 2.5, the latest model in its Composer line. Composer 2.5 scored 62 on our Coding Agent Index, a 14 point gain over Composer 2 (48). This puts it in third place of our tested agents, behind only Claude Opus 4.7 (max) in Claude Code (66) and GPT-5.5 (xhigh reasoning) in Codex (65). These cost $4.10 and $4.82 per task respectively, ~10x the cost of Composer 2.5 Fast ($0.44) and ~60x the cost of Composer 2.5 standard ($0.07). Key results for Composer 2.5 in Cursor CLI: ➤ Cost-quality Pareto frontier: At $0.07 (standard) and $0.44 (Fast) per task, Composer 2.5 is cheaper than every other agent scoring above 60 on the Index. Medium-effort peers cost $1.24–$2.21 per task; higher-effort variants land 3-4 points above at $4.10–$4.82 ➤ Per-benchmark gains vs Composer 2: +35 points on SWE-Bench-Pro-Hard-AA (12% → 47%), +2 points on Terminal-Bench v2 (64% → 66%), and +3 points on SWE-Atlas-QnA (69% → 72%). At 47%, Composer 2.5's score on SWE-Bench-Pro-Hard-AA is comparable to Claude Opus 4.7 (max) in Claude Code ➤ Among the fastest coding agents: Composer 2.5 Fast runs at an average wall time of 6.7 minutes per task, the third-fastest agent on the Artificial Analysis Coding Agent Index, behind only Claude Opus 4.7 (medium) in Claude Code (5.8m) and GPT-5.5 (medium) in Cursor CLI (6.2m) ➤ Fast mode enables better responsiveness at 6x pricing: Fast runs 30% faster than standard Composer 2.5, but is ~6x the cost per task ($0.44 vs $0.07). Token pricing is 6x higher for Fast: $3.00/$15.00 vs $0.50/$2.50 per million input/output tokens Model details: ➤ Base model: Continued training on @Kimi_Moonshot's open weights Kimi K2.5 as with Composer 2, with Cursor reporting ~85% of total compute from its own additional training and reinforcement learning ➤ Pricing: $0.50/$2.50 per million input/output tokens for the standard variant; $3.00/$15.00 for the Fast variant (the default in Cursor) ➤ Available exclusively in Cursor: both Cursor IDE and Cursor CLI, an externally accessible API is not available Congratulations @cursor_ai and @mntruell on the impressive release!

    推荐理由:推文用每任务成本对比 Composer 2.5 与两个更高分编码智能体,读者可据此权衡编码任务上的性能与花费。

  2. Tomer Tunguz82

    SpaceX 递交 S-1,披露 Starlink、发射与 AI 三大业务数据

    SpaceX 递交 S-1,披露 2025 年 187 亿美元合并营收与 66 亿美元调整后 EBITDA。文件把公司分为 Space、Starlink 和 AI 三个分部,Starlink 贡献 61% 营收、2025 年运营利润 44 亿美元,AI 分部当年投入 64 亿美元建设 COLOSSUS 数据中心并训练 Grok。

    推荐理由:S-1 数据把 SpaceX 拆成卫星、发射与 AI 三块业务,读者可借此比较 AI 算力投入与收入回报的差距。

  3. @AYi_AInotes68

    阿易 AI Notes 引用泄露音频称,扎克伯格在 4 月 30 日全员会上表示 Meta 正用员工的键盘、鼠标、屏幕数据训练 AI,认为员工平均智力高于外包,可更快提升 Llama 的编码能力。作者指出 20 天后 8000 名员工收到裁员邮件,并批评这是把员工当免费高质量训练数据、用完就裁的做法。

    引用More Perfect Union (@MorePerfectUS)@MorePerfectUS

    LEAKED AUDIO: In an all-hands meeting on April 30, Mark Zuckerberg tells employees that he's training AI on them ahead of mass layoffs. "The AI models learn from watching really smart people do things... The average intelligence of the people who are at this company is significantly higher than the average set of people that you can get to do tasks. So if we're trying to teach the models coding, for example, then having people internally build tools or solve tasks that help teach the model how to code, we think is going to dramatically increase our model's coding ability faster than what others in the industry have the capability to do, who don't have thousands and thousands of extremely strong engineers at their company." Video

    推荐理由:引用泄露的全员会音频,呈现 Meta 用员工数据训练 AI 与随后裁员的关联叙事,可借此了解事件背景。

5月20日周三
  1. @berryxia73

    Google 发布 Gemini 3.5 Flash,Artificial Analysis 测试显示其 Intelligence Index 为 55 分,比 Gemini 3 Flash 高 9 分,超过 Grok 4.3 和 Claude Sonnet 4.6,输出速度超 280 tokens/s,比上一代快 70%,幻觉率从 92% 降到 61%。

    引用Berryxia.AI (@berryxia)@berryxia

    兄弟们! 今天已经可以在ZenMux上免费体验Gemini 3.5 Flash 了! 我第一时间用它跑了那个经典的「AI模型递归二叉树生长测试」. 同一个 Prompt ,不同模型画出的树形态完全不一样。(见视频-Prompt见评论区) Gemini 3.5 Flash 从输入提示词到生成完整 HTML 动画网页(树干慢慢长出、分支递归展开、最后随风摇摆),全程只用了 77.56 秒! 整体效果非常惊艳:树形态自然优雅、生长动画丝滑、视频和内容呈现都顶级! 熟悉的老朋友都知道,ZenMux 每次新模型都是 ZeroDelay 首发. Google I/O 2026 今天刚发布,现在立刻就能通过 API 调用! 还有免费额度可以白嫖~ 速度是真的没话说,还完美保留了旗舰级模型的能力。 专为 Agent 设计,在 MCP Atlas、Toolathlon、Finance Agent 等多项榜单直接拿下第一! 多模态理解也极强:MMMU-Pro 83.6%、CharXiv Reasoning 84.2%,全面超越上一代 Gemini 3.1 Pro。 完全兼容主流 API 格式,无需改动现有工具链。 支持按量计费 + Builder 套餐。 👇 直接体验 正式版 → zenmux.ai/google/gemini-3.5-… 免费试用 → zenmux.ai/google/gemini-3.5-… Video

    推荐理由:原文用基准与定价的对比说明 Flash 系列的定位变化,读者可据此重新评估轻量模型的成本预期。

  2. @AYi_AInotes70

    Google 在 I/O 上发布 Gemini 3.5 Flash,称其智能与顶级模型相当但输出速度是其他前沿模型的 4 倍,并当天面向所有人开放。作者认为配套的 Antigravity 平台提供桌面端、CLI 和 SDK 全栈开放,目标是做 Agent 时代的 AWS,而 Spark 个人 Agent 只是示范。

    引用Sundar Pichai (@sundarpichai)@sundarpichai

    Just off stage at #GoogleIO, some highlights from this morning 🧵 Gemini 3.5 Flash is available today for everyone in @antigravity and across our products and APIs. Compared to 3.1 Pro, 3.5 Flash is better across almost all benchmarks with huge progress in coding. It’s also comparable to the best models but very fast (4x faster tokens/ second than other frontier models). And when looking at the intelligence versus output speed, it’s in a league of its own in the top right quadrant.

    推荐理由:把胜负手从模型智力转向智能乘速度乘可部署性,作者给出 Google Agent 基础设施的另一种解读。

5月19日周二
  1. @berryxia68

    Anthropic 宣布收购 SDK 与 MCP server 平台 Stainless,该平台自 Anthropic API 早期起就为其生成几乎全部 SDK。作者认为这不只是技术补全,未来 SDK 形态、MCP 协议走向和开发者必须接受的默认行为都会嵌入 Anthropic 自己的产品哲学与安全策略。他由此担心开发者可用的工具链会越来越窄,只剩一种选择。

    引用Anthropic (@AnthropicAI)@AnthropicAI

    Anthropic is acquiring @stainlessapi, an SDK and MCP server platform that has powered every Anthropic SDK since the earliest days of our API. Read more: anthropic.com/news/anthropic…

    推荐理由:作者把 SDK 与 MCP 工具链的归属变化作为切口,讨论这次收购会如何影响开发者的选择空间。

5月16日周六
  1. @op741865

    作者补发飞书 CLI 的 GitHub 地址(github.com/larksuite/cli)并推荐没装的人试试。被引用的分析提到,该 CLI 于 3 月 28 号开源,一个多月达 10000 Star,期间发布 32 个版本、385 个提交,采用面向日常任务、标准 API 和兜底 API 的三层设计,并配套 Skills 作为 Agent 调用说明书。

    引用歸藏(guizang.ai) (@op7418)@op7418

    飞书 CLI 牛皮啊,发布一个月多点就达到 10000 Star 了! 说明用户和市场相当认可这个动作 最近我们可以发现,越来越多的传统办公产品开始发布 CLI 和 Agent。 AI 时代的 SaaS 软件可能得换个做法了:UI 只是最基本的,接下来还要竞争对 Agent 的适配程度以及覆盖率。在这块,我觉得飞书走得相当靠前。 作为一个 IM 软件,飞书在 AI 时代去做这种开放自己所有能力的 CLI 工具,其实是一种非常不传统互联网的尝试。 这对于之前的互联网产品逻辑和经验来说,是一个非常不应该做的决定。 因为他们这个 CLI 几乎可以控制飞书的所有能力:你可以完全不跟飞书的传统 UI 去交互。只跟 CLI 交互,也可以完成飞书上所有的工作。 传统的 IM 办公软件通常非常复杂,入门门槛相对较高。无论从产品逻辑、UI 设计还是交互设计的角度来看,都没有办法太好地消解这种复杂性。 但是 CLI 工具交付给 Agent 以后,就可以快速消解这种复杂性。用户只需要进行对话,这是非常本能的行为,不需要在繁杂的层级列表 UI 里去寻找功能入口。 我拉了一下数据,他们迭代效率也非常恐怖,它们是 3 月 28 号开源的,一个多月发了 32 个版本、385 个提交。 这说明飞书对这块是非常重视的,投入的人力和精力也非常大。 他们在 CLI 本身的设计上也考虑得非常多,下了很多功夫。主要分为三层: 面向日常任务的快捷命令、开放平台对应的标准 API、兜底的 API 调用。 因为人和 Agent 都不喜欢从 2500 个 API 里去寻找参数,但又需要把这些能力暴露出来,所以他们采用了这种分层的形式。 即使做了分层设计,CLI 本身的内容和 API 依然非常多。所以他们把 CLI 作为工具本身,同时做了很多 Skills 用来充当 CLI 的说明书。 Agent 可以分层、分类型地了解应该如何调用这些 CLI 及其命令。 此外,他们在对 Agent 友好的命令包装上做了很多工作,例如: (a) 内置了 Dry Run (b) 结构化输出 (c) 身份选择、权限检查与风险等级评估 (d) 允许 Agent 在发消息前预览请求 (e) 建立了输出格式的“契约”:将成功或失败的结果、原因以及风险提示都放在结构化数据里。 这样如果出错了,AI 可以非常清楚地进行调试和修改,而不是盲目猜测。 其实现在你如果要创业或者做自己的 Agent,就不需要非得写一个界面。 飞书 CLI 加上 Agent 框架可以完成所有的 Agent 产品常见的操作: 你的聊天界面就是你的 Agent 聊天界面; 你的数据库就是飞书多维表格和文档; 你的用户就是把你拉到组织里的群成员;

    推荐理由:飞书 CLI 开源一个多月获 10000 Star,其分层命令与 Skills 设计为 Agent 调用办公软件能力提供了参考。