这是 Gemini 4 Argon 基准测试的预览。 今天开始向网络防御者推出,并尽快向所有人开放。 很高兴看到所有这些进展,迫不及待想让你们都用上!
#模型发布
#模型发布
今日 22 条
Ammaar Reshi@ammaarAI 评分4141
Google AI@GoogleAIAI 评分5353
Google DeepMind@GoogleDeepMindAI 评分3939推出 Gemini 4 Argon——我们的全新前沿模型。 它专为编码、企业知识工作和网络安全防御等复杂工作流打造——今天起通过我们的 Fairwind Program 向一批受信任的测试者开放。

Google DeepMind精选AI 评分7474 Google DeepMind 发布 Gemini 4 Argon,输出上限扩至 1M tokens
Google DeepMind 发布前沿模型 Gemini 4 Argon,先向 Fairwind Program 的可信网络防御者开放,后续将面向开发者、企业和消费者推出。
推荐理由:官方博客给出定价、输出上限和多个基准成绩,可帮读者评估该模型在编码与安全防御场景的实际定位。
Aravind Srinivas@AravSrinivasAI 评分5353引用Perplexity@perplexity_aiWe built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage
Perplexity@perplexity_aiAI 评分4646
MiniMax (official)@MiniMax_AIAI 评分3434引用Creatify Labs@Creatify_LabsIntroducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.
Demis Hassabis@demishassabis精选AI 评分6666引用Pushmeet Kohli@pushmeetVery happy to announce that our team @GoogleDeepmind has pushed the boundaries of generative biology, achieving the successful synthesis of AI-designed proteins that are both functional and watermarked. This proof-of-concept watermarking of the building blocks of life is enabled by SynthID Bio, our new protein watermarking method. It is designed to safeguard the new era of AI-powered generative biology and strengthen global biosecurity. You can read my thoughts here on why watermarking AI-designed proteins is an important research breakthrough: https://x.com/pushmeet/status/2105314763148321102
推荐理由:AI 设计蛋白质首次实现功能与水印兼具并发表于 Nature,同时开源工具,读者可关注生物安全水印路线的实际落地。
Ant Ling@AntLingAGIAI 评分4242
Aidan Gomez@aidangomezAI 评分4949Nils Reimers 和团队做出来的新模型有点疯狂。极强的可扩展性、SOTA 准确率,而且一如既往完全可私有化部署,并可在 Model Vault 上使用!
引用Cohere@cohereIntroducing Cohere Embed 5: our new state-of-the-art family of embeddings models. Get frontier capabilities with Embed 5 Pro or low-latency performance with Embed 5 Fast.
Cohere@cohereAI 评分4444推出 Cohere Embed 5:我们全新的 SOTA 嵌入模型系列。 用 Embed 5 Pro 获得前沿能力,或用 Embed 5 Fast 获得低延迟性能。

Qwen@Alibaba_QwenAI 评分3737Qwen3.8-27B 现已通过 @nebiustf 开放使用。无论你是在构建智能体还是做深度研究,这个 27B 稠密模型都已为你的多步骤工作流准备就绪!🥳
引用Nebius Token Factory@nebiustfQwen3.8-27B is now live on Nebius Token Factory. A compact 27B dense model for coding, research, and agent workflows, with a focus on planning and completing tasks across multiple steps. Start building: https://tokenfactory.nebius.com/endpoints?modals=endpoint-details&model-id=Qwen/Qwen3.8-27B
AK@_akhaliqAI 评分4141
Simon WillisonAI 评分3232 GPT 6.1 Sol:以五分之一价格实现接近 Astra 的智能水平
Simon Willison 于 9 月 29 日发文点评 GPT 6.1 Sol,称其以五分之一的成本提供接近 Astra 的智能水平。该文归入 OpenAI、生成式 AI、LLM 与 GPT 等标签,正文未披露具体参数、benchmark 分数或定价细节。
Yuchen Jin@Yuchenj_UWAI 评分2222Fast:快 2 倍,价格 2 倍。 Ultrafast:快 8 倍,价格 6 倍。 也许我们应该推出一些 Ultra-ultrafast 开源模型端点?

OpenAI@OpenAIAI 评分5555OpenAI 发布 GPT-6.1 Sol,定位为以约五分之一价格提供接近 Astra 水平的智能。官方称其为当前同性能下最具成本效率的模型。

Microsoft Research@MSFTResearchAI 评分3434
Claude Platform release notesAI 评分4141 Anthropic 宣布弃用 Claude Sonnet 4.5,API 将于 11 月 30 日退役
Anthropic 宣布弃用 Claude Sonnet 4.5(claude-sonnet-4-5-20250929),其在 Claude API 上的退役时间定于 2026 年 11 月 30 日。官方建议用户迁移至 Claude Sonnet 5.5,更多细节可查阅模型弃用说明。
Hugging Face BlogAI 评分5959 NVIDIA 发布开源表格基础模型 Kumo Tabular
NVIDIA 发布开源表格基础模型 Kumo Tabular,给定带标签表格后单次前向传播即可完成分类和回归预测,无需训练、调参或特征工程。
OpenAI News精选AI 评分7070 OpenAI 发布 GPT-6.1 Sol,以 Astra 五分之一的价格提供近 Astra 智能水平
OpenAI 发布 GPT-6.1 Sol,定位为接近 Astra 智能水平的模型,主打编码、计算机使用和专业工作场景,价格为 Astra 标准 API 输入和输出 token 价格的五分之一。
推荐理由:原文明确了模型定位与五分之一的价格对比,读者可以据此评估在不同工作负载下替换现有模型API的成本空间。
inclusionAI Hugging Face modelsAI 评分5656 inclusionAI 发布开源全模态模型 Ming-flash-omni 2.0
inclusionAI 在 Hugging Face 发布开源全模态模型 Ming-flash-omni 2.0,基于 Ling-2.0 MoE 架构,总参数 100B、激活 6B,称在开源全模态 MLLM 中达到 SOTA。
Latent SpaceAI 评分3939 Claude Opus 5.5 发布:SimpleBench 88.4% 登顶,擅长讲解视频
Claude Opus 5.5 本周发布,以 88.4% 登顶 SimpleBench,并被 Anthropic 评为迄今最强视觉模型,成本比 Fable 5.1 低约 60%。在 Terminal-Bench-Science 上,其得分从低推理投入的 24% 升至 xhigh 的 62%,max 档反降至 59%。社区反馈称 200 美元的 Claude Code 套餐已胜过 Codex。
Simon WillisonAI 评分7474 Anthropic 发布 Claude Sonnet 5.5,并成为 claude.ai 免费层模型
Anthropic 发布 Claude Sonnet 5.5,官方称速度提升 30% 以上,多数工作成本降低最多 30%,定价与 Sonnet 5 相同但在各项基准上更强。
Claude@claudeaiAI 评分2929一组 Claude Sonnet 5.5 的早期实验。 @_re_pete 用 Sonnet 5 与 Sonnet 5.5 制作的秋叶模拟器对比。

Thariq@trq212精选AI 评分7272引用Claude@claudeaiIntroducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
推荐理由:作者结合 token 成本这一常见顾虑,指出 Sonnet 5.5 与 Opus 5.5 让更高层智能更易负担,适合在构建工作流时选用。
Charlie Holtz@charlieholtz精选AI 评分6666引用Claude@claudeaiIntroducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
ClaudeDevs@ClaudeDevs精选AI 评分7474
引用Claude@claudeaiIntroducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
推荐理由:原文给出 Sonnet 5.5 相对 Sonnet 5 的速度、成本与适用任务,开发者可据此判断是否切换日常 Claude Code 用法。
Boris Cherny@bchernyAI 评分6161
引用Claude@claudeaiIntroducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
Dongxi 东锡 NLP@dongxi_nlp精选AI 评分7575引用Claude@claudeaiIntroducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
推荐理由:官方发布说明给出了相对 Sonnet 5 的速度提升与降价幅度,读者可据此权衡换用成本。
Anthropic@AnthropicAI精选AI 评分7676引用Claude@claudeaiIntroducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
推荐理由:Anthropic 官宣 Claude Sonnet 5.5 上线,直接给出比 Sonnet 5 快 30%、多数工作成本低 30% 的关键变化。
Claude@claudeai精选AI 评分7575Anthropic 发布 Claude Sonnet 5.5,是 Claude 5.5 家族的第二款模型。相比 Sonnet 5 明显升级,运行速度快 30% 以上,多数工作成本最高降低 30%。

推荐理由:官方给出较 Sonnet 5 的提速和降价幅度,读者可据此评估升级或切换成本。
Kling AI@Kling_aiAI 评分5555
Hugging Face Blog精选AI 评分7575 H Company 发布 Holo4 系列通用计算机操作智能体模型
H Company 发布 Holo4 智能体模型系列,包含 27B dense 和 35B-A3B MoE 两个尺寸,并附带基于 Nemotron 3 Nano Omni 后训练的 Holotron4 Nano。
推荐理由:官方发布给出了跨 GUI、代码、MCP 和 API 四类接口的统一智能体模型,附基准分数和完整轨迹数据,适合评估开源方案与闭源模型的成本差距。
MiniMax (official)@MiniMax_AIAI 评分4747引用MiniMax_Agent@MiniMaxAgentMiniMax's latest text model, M3.1-Flash-Preview, debuts today on MiniMax Code. Built for everyday development, it's fast, reliable, and ready for real work, from quick bug fixes to full features.
Claude Platform release notes精选AI 评分7878 Anthropic 发布 Claude Sonnet 5.5(claude-sonnet-5-5)并给出迁移指南
Anthropic 发布 Claude Sonnet 5.5(claude-sonnet-5-5),可在 Claude API、Amazon Bedrock、Claude Platform on AWS、Google Cloud 和 Microsoft Foundry 使用。
推荐理由:官方列出了五类从 Claude Sonnet 5 升级会破坏的代码场景和迁移指引,方便开发者提前排查自己的 API 调用。
Anthropic Newsroom精选AI 评分8585 Anthropic 发布 Claude Sonnet 5.5:速度提升 30% 以上、单任务成本最多降 30%
Anthropic 发布 Claude Sonnet 5.5,为 Claude 5.5 家族第二款模型,生成速度比 Sonnet 5 快 30% 以上,测试中单任务成本最多低 30%。
推荐理由:官方发布给出了与 Sonnet 5 和 Opus 5.5 的多项基准对比、定价和成本数据,读者可以据此判断它在自己场景下是否值得迁移。
ViggleAI@ViggleAIAI 评分4444引用Yun Chen@t_muxviggle-turbo for Qwen-Image-2.1 isn't just faster — for most prompts, it's just as good as base.
Claude@claudeai精选AI 评分6666引用Ryan Sael@RyanSaelI asked Opus 5.5 to explain camera focus by building an interactive lens lab Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost https://lens.lab.sael.net Move the focus ring and you can see the glass elements shift the sharp plane through the scene
推荐理由:官方汇总用户用 Opus 5.5 探索的成果,引用案例给出了单次生成时长与成本,可作实际使用参考。
karminski-牙医@karminski3AI 评分5959
引用Meituan LongCat@Meituan_LongCatLongCat-2.5-Preview is now live. 1.6T parameters. ~48B active. A 1M-token context window. Natively multimodal. Built to take on long-horizon tasks. From terminals and browsers to GUIs, spreadsheets, and design tools. Try it now: 🚀 API: https://longcat.ai/platform/ 💬 Chat: https://longcat.ai/chat/
Latent SpaceAI 评分3636 Latent Space 宣布 AINews v3 改版并重启赞助合作
Latent Space 计划下周推出 AINews v3,将已运营 3 年、订阅超 20 万的 AINews 与 Latent Space Discord 合并,并探索迁移至 Beehiiv 和新首页。