跳到正文

#部署/工程

今日 19 条
9月22日周二
  1. OpenRouter Announcements61

    OpenRouter 解析 NVIDIA Nemotron 3.5 Lightning 如何承担智能体高频执行调用

    OpenRouter 发文解析 NVIDIA 的 Nemotron 3.5 Lightning,这是一个 30B 总参数、每 token 激活约 3B 的 MoE 开源权重模型,定位于智能体工作流中的高频执行步骤。

    推荐理由:原文梳理了模型规格、与 Ultra 的分工、各端点差异和路由方法,可帮助读者判断何时用它替代大模型执行高频调用。

  2. Google Developers Blog42

    Google Colab 纳入 Google AI 订阅计划

    Google AI 订阅用户现可直接获得 Colab 高级权益,包括优先使用更快的加速器和更强算力的机器。Google AI Ultra 订阅者还可解锁不间断后台执行和 Premium GPU 访问,长时间训练无需保持浏览器标签页开启。该权益将在未来几周内向 Colab 支持国家的 Google AI 订阅用户逐步推送,已有 Colab 订阅不受影响。

  3. OpenRouter Announcements71

    OpenRouter 推出 Batch API,批量推理约半价

    OpenRouter 推出 Batch API,提交批量请求后供应商可在 24 小时窗口内完成,per-token 价格通常为常规价的 50% 或更低,已支持 70 多个模型。

    推荐理由:官方给出了实测的批次完成时长分布和提交时段差异,便于读者判断批量推理能否接入现有工作流。

9月21日周一
  1. Qwen47

    感谢 @sgl_project 的 day-0 支持!🙌 SGLang-Diffusion 现已支持 Qwen-Image-2.1:文生图、多图编辑,以及透明 RGBA 输出。快来试试!🎨

    引用SGLang@sgl_project

    Day-0 support for @Alibaba_Qwen’s Qwen-Image 2.1 is here in SGLang-Diffusion! 🖥️ Native precision on a single RTX 4090 24GB with CPU offload - 1024×1024 generation in 18.7s and image editing in 21.7s with 22.7 GiB peak GPU memory during requests. - On an RTX PRO 6000 96GB: 8.0s generation and 9.6s editing. 🎨 Text-to-image, multi-image editing, and transparent RGBA output—all with one checkpoint. ⚡ Native inference with TP/SP, LoRA, and OpenAI-compatible APIs. 40 denoising steps, one image per request, warmed HTTP latency including PNG output. No quantization. Cookbook and GPU-specific commands below 👇

  2. Hugging Face Blog57

    Hugging Face tokenizers v1 发布候选,编码速度较 v0.23 提升最多 30 倍

    Hugging Face 发布 tokenizers v1 候选版本,单线程编码速度在十个模型族上比 v0.23 快 3 到 30 倍(Apple M4 Max,低端 t5-base,高端 gpt2),八线程扩展达线性 76%,且 token ID 与旧版完全一致。主要改动包括用 SIMD 位流切分替代正则、线程本地词缓存、免分配的合并循环和按批模型调用,候选版已在 crates.io 提供。

9月20日周日
9月18日周五
  1. MiniMax (official)34

    Nunchux AI 与 MIT、CMU、UC Berkeley、斯坦福及 NVIDIA 研究者合作推出 VC-Attention,为 MiniMax-H3 带来免训练低比特注意力加速,在 B200 上比 FlashAttention-4 快 1.6×、B300 上快 1.5×,保真度优于 SageAttention2。

    引用Nunchux AI@NunchuxAI

    Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.

9月17日周四
  1. jietang61

    唐杰称,由 GLM-5.3 驱动的基础设施智能体用两周时间让 GLM-5.3-Flash 从首次在国内加速器上运行到承接全部生产流量,端到端吞吐提升 3.2 倍。

    引用Z.ai@Zai_org

    We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://z.ai/blog/glm-built-its-inference-infrastructure

  2. GitHub Blog · AI & ML78

    GitHub 用 Copilot 将 Copilot agent runtime 迁移到 Rust

    GitHub 用 Copilot 智能体将 Copilot agent runtime 从 TypeScript 完整重写为 832,378 行生产 Rust,共 128 个 PR,于 8 月 21 日完成。

    推荐理由:作者以第一手移植经历拆解了 AI 智能体团队完成大规模重写的具体策略、会话数据和教训,方法细节对类似工程迁移有直接参考价值。

  3. OpenRouter Announcements67

    OpenRouter 教程:用 TypeScript SDK 构建可靠的工具调用 Agent 循环

    OpenRouter 发布教程,演示如何用其 TypeScript SDK 从零构建工具调用 Agent 循环,示例使用本地天气数据可独立运行。教程覆盖停止条件设计、按 toolCallId 返回结果、指纹计数拦截重复调用、models 参数实现有序模型回退,以及 Auto Exacto 默认按工具调用成功率重排提供商;并说明 MCP 只改变工具执行位置,循环控制仍是必要的。

    推荐理由:教程给出了完整的停止条件、重复检测和模型回退实现,可用作自建 tool-calling 循环的参考骨架。

9月16日周三
  1. ByteByteGo52

    LLM 如何在海量文档中找到关键信息:RAG 检索链路详解

    ByteByteGo 详解 LLM 应用中的检索问题,以员工询问航班取消后酒店报销为例,说明 RAG 如何从数千份文档中找到正确且仍然有效的政策条款。文章覆盖 chunking 粒度权衡、嵌入向量的语义匹配、余弦相似度等度量选择、Flat/IVF/HNSW 索引的速度与召回权衡、元数据过滤的先筛后筛取舍,以及政策更新时的版本切换和混合检索加 reranking 的最后优化步骤。

  2. Hugging Face Blog63

    IBM Research 为 ALTK-Evolve 推出 Consistency Analyzer,将智能体一致性缺口从 24.4pp 降到 12.0pp

    IBM Research 在 ALTK-Evolve 中新增 Consistency Analyzer 和一致性指南,用于衡量并改善智能体重复执行同一任务时的稳定性。

    推荐理由:原文给出 Pass^k 与 Mean@k 的区别和一种只靠单条轨迹、无需标注的稳定性诊断方法,可将智能体一致性缺口缩小约一半。

9月15日周二
  1. ByteByteGo59

    ByteByteGo 图解 LLM 的记忆机制:模型本身不会记住对话

    ByteByteGo 发文解释 LLM 为何在新会话中忘记此前对话:模型本身没有持久记忆,记忆错觉由围绕模型的应用通过重建上下文实现。文章区分训练记忆、上下文窗口工作记忆和外部持久记忆,分析上下文填满后的处理方式与长对话成本,并介绍滑窗、对话摘要、结构化实体抽取、向量库检索和长期用户画像等扩展记忆的技术。

  2. Claude Blog69

    Claude for Small Business 新增 43 个工作流与 27 个集成,并重启 SMB 巡回培训

    Anthropic 为 Claude for Small Business 推出 43 个工作流和 27 个新集成,覆盖 Shopify、Salesforce、Stripe、Zoom、Xero、Zapier 等工具,产品自 5 月上线以来安装量超过 900,000 次。

    推荐理由:官方公布了新增工作流与集成数量,并给出多个客户的可验证使用结果,读者可据此评估这类 Agent 产品对小型业务的落地形态。

9月14日周一