

关注 AI 研究者、开发者与机构的动态


Gemini 3.5 Transcribe 是我们最新的语音转文字模型,可实现精准且智能的转录。🧵
我们将在未来几天添加更多有用信息,比如所属机构、引用等。如果你有任何想法,欢迎告诉我们。 https://t.co/5oeUXtIpyw
过去 5 年多我一直在整理 AI 论文。 现在我们把所有论文索引到了一个地方。 你可以按主题发现有趣的论文,还能和它们对话。尽情享用!https://t.co/Rxz4Sv8nGQ
一款高性能 125B 模型现在仅需 75GB RAM 即可在本地运行!感谢 @UnslothAI 的首日支持。🥳 https://t.co/za7d6TIluk


在 LiveKit 中使用 Grok Voice 模型,构建实用的语音智能体,全面支持 ZDR
We built a patient intake agent with @SpaceXAI that runs end to end on Grok voice models. Grok STT → Grok 4.3 → Grok TTS, cascaded through LiveKit Inference. Three model strings in one AgentSession. No separate API key, no separate billing, ZDR on every hop. Talk to it: https://xai.livekit.space
彭博确认在 OpenRouter 匿名屠榜的 Ox Alpha 是智谱 GLM 系列的新迭代,今晚将直接开源放权重。
牛来大模型刚被彭博社破案了! 在 OpenRouter 上悄悄屠榜、调用量超过 DeepSeek 一倍的神秘模型 Ox Alpha, 背后竟然是智谱,更绝的是官方确认今晚直接开源权重, 狂刷了数十万亿 token、在真实 Coding 里跑出 80% 峰值,智谱这次是真的要杀疯了, 前阵子大家都在猜这个匿名模型到底是哪家大厂的马甲, 因为它在排位赛上狂刷了数十万亿 token, 开发者用脚投票把它送上了第一, 它真正狠的地方,是把 Flash 级别的速度和成本,做出了中上游前沿模型的实战能力: 1M 超长上下文、原生支持图片和视频多模态, 而且是推理优先的架构,写代码时会先做规划再调工具 社区独立跑测试,在 10 个高难度真实 Coding 任务里拿下了 80% 的过关率, 全量 DeepSWE 跑出 64.6%, 直接咬住了 Claude Opus 4.8 和 Gemini 3.7 Flash 的身位 它不是那种打虚空跑分的绝对 SOTA, 但一个成本极低、速度极快、能看懂视频还能写复杂代码的 Flash 模型, 今晚一旦把权重放出来,无论是本地部署还是第三方 API,生态杀伤力完全不可同日而语 刚发完开源顶流 GLM-5.3,转头就把这个高频 Agent 生产力核弹开源出来, 大模型的竞争,终于从比拼谁的参数更大,彻底转向谁能让开发者用最低成本把活干完了啊
推荐理由:匿名屠榜模型被确认为智谱 GLM 新迭代并将开源权重,读者可了解它对第三方 API 与本地部署成本的影响。



Finally, Z .ai revealed that Ox Alpha was actually GLM-5.3-Flash. So that means over the last few days all those 100 tn tokens/day of stealth traffic capacity was running on Chinese AI chips, with tens of thousands of domestic accelerators behind the service. not an NVIDIA GPU cluster. 5.3-Flash beats GLM-5.2 at one-tenth the price with only 18B active parameters. GLM-5.3-Flash. has 320B params in total, but only 18B are active during inference. It also uses 45 layers instead of GLM-4.5's 92, cutting the amount of work required for each token. The benchmark jumps are large too: against GLM-5.2, DeepSWE rises from 46.2 to 63.4 and AutomationBench from 26.2 to 48.8. There is an architectural change as well, that cuts attention compute 3x and per-layer KV cache 4.4x versus GLM-5.3. The novelty is mainly in the combination: GLM-5.3-Flash uses linear attention for cheap state tracking, then sparse attention with a lightweight indexer to retrieve only the distant context worth revisiting, instead of repeatedly attending across the full 1M-token window. They also introduces IndexPool, which compresses four indexer key vectors into one, and says the combined design
阿里云发布开放权重的 Qwen3.8-Flash,这是一款多模态 MoE 模型,也是 Qwen4 架构的早期预览。
推荐理由:作为 Qwen4 架构的前置预览,它给出混合注意力、N-gram 嵌入等设计细节,并可与 Qwen3.7-Plus 的成本和成绩对照。


感谢 @ArtificialAnlys 在同一套手机端量化版本上评估智能水平与真机推理。 查看完整基准与方法论: https://t.co/vHGpiekGRk
$visualize https://t.co/rItCXe9sgm https://t.co/F5U0PlDvRp
Brain is our self-improving memory system for Perplexity Computer. It compiles sessions, files, and sources into a structured knowledge wiki. New evals build on our initial results, improving correctness by 9.3 points, currentness by 8.0, and recall by 8.9 with 15% fewer tokens. https://t.co/mDMVWt2xzS
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG
推荐理由:训练成本降到 Qwen3.7-Plus 的 1/9 且官方称全面超越,读者可借此观察 Qwen4 新架构的取舍。
Z.ai 披露此前匿名上线的模型 Ox Alpha 实为 GLM-5.3-Flash,作者转述的报道称该服务由数万个国产加速器支撑,而非 NVIDIA GPU 集群。
OxAlpha is a new iteration of GLM, from China’s Z .ai And it will change how you run long running agent fast. --- bloomberg .com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek https://t.co/ZOTlnBgOsy
推荐理由:原文确认 Ox Alpha 即 GLM-5.3-Flash,并给出参数规模、架构改动与基准对比,便于对照其相对 GLM-5.2 的位置。
微信开源了 WeMM-Embedding 多模态嵌入模型,包含 2B、4B、9B 三个版本,其中 9B 拿下双榜第一,2B 版本的跑分超过上一代 8B 开源模型。
Qwen3.8-Flash-Next 与 GLM-5.3-Flash 两款开放权重模型发布,均为 MoE 架构,每 token 分别激活 6B 和 18B 参数。


推荐理由:推文列出两款开放权重模型的参数规模与多项基准结果,可供读者对比其与前沿闭源模型的差距。


Brain 是一套全面的记忆系统,包含三个关键组件:持久化记忆存储、回答查询的前台智能体,以及更新和改进记忆的后台智能体。https://t.co/BidkEvm9nz
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG
推荐理由:发布信息列出了 Qwen3.8-Flash 的参数量、激活规模与定价,可作为判断高效多模态 MoE 路线的具体参照。