Strands Agents 与 LeRobot 如何从 Hugging Face Hub 部署到机器人硬件
AWS 以 Apache 2.0 开源的 Strands Robots 把 LeRobot 的硬件抽象、仿真和策略推理封装成 AgentTools,让单个 Strands 智能体从 Hugging Face Hub 数据集走到真实机器人。
推荐理由:把从 Hub 数据集到物理机器人的流程拆成五步,仿真与真机共用同一数据集格式,可作为具身智能体接入的参考。
开源模型、框架与仓库动态:权重开放、社区项目爆火、开源与闭源的力量消长。
当前仅显示精选新闻AWS 以 Apache 2.0 开源的 Strands Robots 把 LeRobot 的硬件抽象、仿真和策略推理封装成 AgentTools,让单个 Strands 智能体从 Hugging Face Hub 数据集走到真实机器人。
推荐理由:把从 Hub 数据集到物理机器人的流程拆成五步,仿真与真机共用同一数据集格式,可作为具身智能体接入的参考。
GLM-5.2 发布,是面向长周期任务的旗舰模型,首次在 1M token 上下文上提供该能力,并采用 MIT 开源协议。相比 GLM-5.1,它在 Terminal-Bench 2.1 上从 63.5 提升到 81.0,在 SWE-bench Pro 上从 58.4 提升到 62.1。
推荐理由:GLM-5.2 将上下文扩展到 1M token 并以 MIT 协议开源,读者可对照它与 Claude Opus 4.8、GPT-5.5 在长周期基准上的差距。
智谱上线并开源 GLM-5.2,在全球百万用户参与盲测的前端开发评估系统 Code Arena 上取得全球可用模型第一;摩尔线程同日宣布在 AI 训推一体全功能 GPU 智算卡 MTT S5000 上完成 Day-0 极速适配。
推荐理由:材料给出国产 GPU 对 GLM-5.2 的 Day-0 适配细节,可观察长上下文推理的软硬件协同路径。
智谱发布 GLM-5.2,以 MIT 协议完全开源,参数 753B,支持 1M token 上下文窗口,主打长程任务能力。文中称在相近 token 消耗下其能力介于 Opus 4.7 与 Opus 4.8 之间;FrontierSWE 得分 74.4,落后 Opus 4.8 约 1%,超过 GPT-5.5 的 72.6。
推荐理由:文中给出 GLM-5.2 在长程编程基准上与 Opus 4.8 的分数对比,可看到开源模型在长任务上的位置。
Tomer Tunguz 梳理了一篇 Hacker News 热帖的 500 多条评论,勾勒出本地编码栈的现状:Qwen 3.6 35B-A3B 以 33% 的提及率居首,27B 变体占 20%,DeepSeek Pro 与 Gemma4 31B 进入前四;智能体框架方面 Pi 以 49% 领先,OpenCode 紧随其后达 45%。
推荐理由:帖子数据勾勒出本地编码模型与智能体工具的占比,并与 Claude 做能力对比,便于判断本地替代的可行边界。
MiniMax 开源了 MiniMax M3 模型权重,并同步发布 MSA(MiniMax Sparse Attention)技术论文。M3 是 MiniMax 的原生多模态旗舰模型,总参数 428B,激活参数 23B,文中称其是第一个从 Step 0 开始做多模态混合训练的开源模型。
推荐理由:M3 从 Step 0 起做多模态混合训练并开源权重,同时公开 MSA 稀疏注意力论文,可了解其长上下文成本设计思路。
inclusionAI 发布 Ming-omni-tts-16.8B-A3B 统一音频生成模型,可在单通道内联合合成语音、环境声与音乐。模型基于自研 12.5Hz 连续 tokenizer 与 Patch-by-Patch 压缩,将 LLM 推理帧率降至 3.1Hz,并内置 100+ 音色与零样本语音设计。
推荐理由:该模型在单通道内联合生成语音、环境声与音乐,并用方言与情感控制基准对比 CosyVoice3 与 Qwen3-TTS。
开发者 Jamieson O'Reilly 用泄露的 Claude Fable 5 系统提示词,通过一行 claude --dangerously-skip-permissions --system-prompt-file CLAUDE-FABLE-5.md 命令将其注入 Opus 4.8,在生成苹果风格落地页时重现了 Fable 5 的风格。
推荐理由:开发者用泄露提示词让 Fable 5 在 Opus 4.8 上重现风格,文章同时梳理亚马逊测试触发政府禁令的内幕。
AllenAI 发布 olmo-eval,一个面向模型开发循环的开源评测工作台,在其 OLMES 评测标准基础上扩展。它把任务、套件与 harness 解耦,支持智能体与多轮评测,并可按需选择直接运行或容器沙箱执行,以控制成本。除整体分数外,它还给出标准误和最小可检测效应,并能将两个模型 checkpoint 逐题对齐比较,以区分真实改进与噪声。代码已在 GitHub 开放。
推荐理由:它把评测从一次性跑分挪进持续训练循环,逐题对比 checkpoint 能看出被平均分掩盖的真实变化。
Meet DiffusionGemma ⚡ Our latest experimental open model (Apache 2.0) that generates text up to 4x faster. Instead of predicting and typing just one word at a time like most language models, it drafts and refines entire blocks of text simultaneously. Here’s how it works 🧵 ↓
推荐理由:Google 开源实验性文本生成模型,采用并行 diffusion 思路,读者可了解其与主流自回归生成在路线上的差异。
Google 发布 Apache 2 许可的开源权重 Gemma 模型 google/diffusiongemma-26B-A4B-it,NVIDIA 目前在其 NIM 云 API 上免费托管该模型。Simon Willison 用该 API 生成图片,返回 2,409 tokens 耗时 4.4 秒,即每秒至少 500 tokens。
推荐理由:作者用 NVIDIA 免费 API 实测这款开源权重 Gemma 模型,给出生成速度的实测参考。
Google DeepMind 发布 Apache 2.0 开源的实验性文本扩散模型 DiffusionGemma,为 26B MoE 架构、推理时激活 3.8B 参数,在专用 GPU 上文本生成最高提速 4 倍,单张 NVIDIA H100 超过 1000 tokens/秒,量化后可放入 18GB 显存。
推荐理由:官方给出了具体吞吐数字、显存占用和适用场景边界,读者可以据此判断扩散文本生成适合哪些本地工作流。
Google DeepMind 发布 Gemma 4 12B,采用无编码器统一架构,视觉与音频输入直接进入 LLM backbone,是该系列首个支持原生音频输入的中等规模模型。
推荐理由:官方说明新模型以无编码器统一架构把性能接近 26B 的多模态能力压到 16GB 内存笔记本可跑,开发者可据此评估端侧部署选择。
OpenEnv 宣布由 Meta-PyTorch、Reflection、Unsloth、Modal、Prime Intellect、Nvidia、Mercor、Fleet AI、Microsoft、Hugging Face 和 RadixArk 组成的委员会共同协调,项目地址迁移至 huggingface/OpenEnv。
推荐理由:OpenEnv 由多家机构组成的委员会共同协调,并明确为 RL 环境的互操作层,读者可了解开源智能体训练的协作与协议设计。
Supabase 宣布完成 5 亿美元的 F 轮融资,公司估值达到 100 亿美元。而在一年前其估值仅为 20 亿美元。
x.com/i/article/206313956911…
推荐理由:Supabase 一年内估值从 20 亿美元升至 100 亿美元,这条信息可看出 AI 时代基础设施公司的融资热度。
NVIDIA 发布 Nemotron 3.5 Content Safety,基于 Google Gemma 3 4B IT 微调,把多模态输入、自定义企业策略与可审计推理链统一到单次推理调用中,并保持 12 种语言显式训练和约 140 种语言的零样本泛化。
推荐理由:相比 Nemotron 3,3.5 版把多模态审核、自定义策略与可审计推理链合并到一次调用中。
Introducing Magenta RealTime 2 (MRT2): the live music model you can play as an instrument. MRT2 offers MIDI and prompt controls, and runs natively on a MacBook with <200ms latency. Open weights. Open source inference engine. Suite of apps and plugins. Hear what it can do and try it out for yourself below 🧵 Video
推荐理由:开放权重与开源推理引擎一并给出,读者可了解实时音乐模型在笔记本上的本地延迟表现。
NVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest US open weights models, Gemma 4 31B (39.2), Nemotron 3 Super (36.0) and gpt-oss-120b (33.3), but behind the Chinese-led open weights frontier (Kimi K2.6 at 53.9). We partnered with @NVIDIA to evaluate this model for intelligence and speed ahead of its public release. These figures use the final NVFP4 weights that NVIDIA recommends for inference, but our tests show minimal intelligence impact compared to BF16 testing, with higher precision resulting in an Artificial Analysis Intelligence Index score of 48.2 vs. the NVFP4 score of 47.7. Key Takeaways: ➤ Nemotron 3 Ultra leads in speed for its intelligence: through BlackBox AI ahead of release, Nemotron 3 Ultra is served at over 400 output tokens per second - this is slightly faster than the typical serving speed of gpt-oss-120b despite being >4X larger, and comes with significantly greater intelligence ➤ Largest Nemotron 3 model so far: with approximately 550 billion total parameters and 55 billion active, Nemotron 3 Ultra is significantly larger than its siblings and is the largest and most intelligent US open weights model release ever ➤ Nemotron 3 Ultra is the leading US open weights model on the Artificial Analysis Intelligence and Agentic Indexes by far, but Gemma 4 31B scores ~1 point higher on the Coding Index (comprised of Terminal-Bench Hard and SciCode)
推荐理由:Artificial Analysis 的评测让读者能横向比较 Nemotron 3 Ultra 与美国及中国开源权重模型的智能与速度表现。



Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models. Video
推荐理由:原文给出 550B 开源模型的权重、训练数据与完整配方,并说明其在长任务智能体上的吞吐表现,读者可据此判断开源前沿模型的可复现程度。