Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸
#端侧
#端侧
今日 8 条
Chubby♨️@kimmonismusAI 评分4545
引用David Stout@Davidstout
Chubby♨️@kimmonismusAI 评分4343
IT HomeAI 评分2727 惠普战 99 锐龙版 AI 移动工作站发布:AMD R9-9955HX3D + RTX PRO 2000,21999 元
惠普战 99 锐龙版 AI 高性能移动工作站于 10 月 1 日发布,10 月 12 日晚开售,顶配 R9 9955HX3D + RTX PRO 2000 + 32GB 内存 + 1TB 硬盘售价 21999 元。
IT HomeAI 评分4242 微软预热本月 Surface 发布会:黄仁勋出席,聚焦本地 AI 与 RTX Spark
微软将于太平洋时间 10 月 7 日举办面向开发者与构建者的线上活动,主题为本地 AI 如何塑造下一代 PC,英伟达 CEO 黄仁勋与微软 CEO 萨提亚·纳德拉等将出席。
TechCrunch · AIAI 评分5555 Satlyt 完成 800 万美元种子轮融资,为卫星提供在轨 AI 运行软件
由前 Google 和 SpaceX 产品经理 Rama Afullo 联合创办的 Satlyt 完成 800 万美元种子轮融资,由 Non Sibi Ventures 领投,为多公司卫星提供运行 AI 模型的软件,对标 VMware 和 Snowflake 的平台模式。
TechCrunch · AIAI 评分4848 Legato 推出 AI 助听眼镜 Legato Frames,起售价 999 美元
听力科技初创公司 Legato 周四宣布,其 AI 助听眼镜 Legato Frames 正式开售,起售价 999 美元,面向轻至中度听力损失成年人。该眼镜将专利助听技术集成于镜腿,AI 系统可区分人声与背景噪音,双扬声器系统在耳旁数英寸处降低 99% 漏音,重 34 克,续航 10-12 小时,支持处方镜片和蓝牙串流。
TechCrunch · AIAI 评分6666 Google 首次将 TPU 送入太空,其研究认为 Starship 需发射约 1,800 次轨道数据中心才可行
Google 的 Project Suncatcher 原型卫星由 SpaceX 火箭发射升空,将验证 TPU 能否在太空运行。
elvis@omarsar0AI 评分4848
引用David Stout@DavidstoutHalf a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸
Simon WillisonAI 评分5858 Simon Willison 用 GPT-6 Astra 构建本地人脸模糊与元数据移除工具 Photo Scrubber
Simon Willison 发布实验性工具 Photo Scrubber,可自动识别照片中人脸并模糊处理,用于分享陌生人照片前保护隐私。工具基于 Google 的 MediaPipe C++ 库,经 @mediapipe/tasks-vision 编译为 WebAssembly,并使用 BlazeFace 人脸检测模型,由他用 GPT-6 Astra 构建产生。
Google Cloud: Databases精选AI 评分6565 Google Cloud 分析创业公司为何需要在前沿 API 之外搭配 Gemma 4 开源模型
Google Cloud 发文主张创业公司采用“复合 AI 栈”,用开源的 Gemma 4 处理边缘执行、高吞吐分流、任务微调和垂直场景,把 Gemini 留给复杂推理。
推荐理由:文章用三个创业案例和四类工作负载说明开源模型与前沿 API 搭配的架构取舍,适合正在做模型选型的团队参考。
Perplexity@perplexity_aiAI 评分4343
Hugging Face BlogAI 评分5959 Liquid AI 发布 LFM2.5-VL-3B 的 DSpark 草稿模型,解码最高提速 3.13x
Liquid AI 为视觉语言模型 LFM2.5-VL-3B 发布实验性 DSpark 草稿模型,通过投机解码在不改变输出质量的前提下加速推理。解码最高提速 3.13x(M5 Max 设备端)和 2.66x(H100),端到端最高提升 2.62x 和 2.27x。
OpenBMB@OpenBMBAI 评分3838


inclusionAI Hugging Face modelsAI 评分5959 inclusionAI 发布 Ling-3.0-tiny:7.9B 总参数、1.3B 激活的混合推理 MoE 模型
inclusionAI 发布轻量混合推理 MoE 模型 Ling-3.0-tiny,总参数 7.9B,每 token 仅激活 1.3B,采用 3:1 KDA-MLA 混合线性架构加 128 专家稀疏 MoE FFN。
Tencent Hy@TencentHunyuanAI 评分4646
Apple Machine Learning ResearchAI 评分3939 Apple 用潜空间蒸馏压缩流式神经音频编码器,2.8 倍压缩下 WER 仅升 1.9%
Apple 研究团队提出潜空间蒸馏方法压缩流式神经音频编码器(tokenizer),以教师模型量化前的逐帧潜表示为监督目标,仅训练学生编码器做回归,并用单层仿射层吸收师生宽度差异。在 2.8 倍压缩下,六组师生配对中有五组无需微调即保持在教师 1.9% 相对 WER 以内,并比同容量独立训练编码器相对提升 3.9%。该方法同时适用于独立预训练和与语言模型联合训练的 tokenizer。
elsewhere articlesAI 评分3333 弋途科技完成近亿元 Pre-B 轮融资,AICAR 正式上线
弋途科技(EXTURING)近日完成近亿元 Pre-B 轮融资,由上海半导体装备材料产业投资基金与 Sands Talk Capital 联合领投,资金将用于加速 AICAR 规模化推广及自研端侧模型落地。公司已服务超 70% 国内头部车企及主流合资品牌,数十款车型量产上车,累计交付达百万级,AICAR 近期正式上线。
Google Developers Blog精选AI 评分6565 Antigravity SDK 支持本地模型,首发接入 Gemma 4 26B A4B 与 LiteRT
Google 宣布 Antigravity SDK 支持本地工作流,首发支持通过 Google AI Edge 的 LiteRT 运行 Gemma 4 26B A4B,可完全离线提供智能体能力,建议机器配备 24GB 以上 VRAM 或统一内存。
推荐理由:原文给出本地运行智能体的具体配置方法和混合编排演示数据,开发者可以直接照此把智能体工作流搬到本地 GPU 上。
OpenBMB@OpenBMBAI 评分5858
Hugging Face BlogAI 评分4848 oMLX 创作者 Jun Kim 加入 Hugging Face,支持 MLX 社区
oMLX 创作者兼维护者 Jun Kim 加入 Hugging Face,全职投入 MLX 社区建设。oMLX 将从副业转为有资金支持的正式项目,继续以 Apache 2.0 开源并由 Jun 领导,目标是加快开发并更好引导贡献者。Hugging Face 计划让 oMLX 成为新想法的试验场,并推动 transformers 模型定义快速转为可被不同引擎使用的 MLX 参考实现。
Hugging Face Blog精选AI 评分7474 transformers 支持直接加载 llama.cpp GGUF 量化模型
Hugging Face 宣布 transformers 支持通过 from_pretrained 加载 GGUF 量化模型,复用 ggml 的 Metal 内核,初始聚焦 Apple Silicon 上的 Qwen3.5 架构。
推荐理由:原文给出加载方式、性能对比数据和方法论,读者可据此判断在 Mac 上用 transformers 跑 GGUF 的可行路径。
ByteByteGoAI 评分5757 如何在廉价硬件上运行大模型:从量化到投机解码的推理优化详解
ByteByteGo 撰文讲解如何在普通硬件上运行大模型推理。文章解释内存容量、带宽和 prefill 与 decoding 两阶段的瓶颈,并逐一介绍量化(8B 模型 16-bit 权重约 16 GB。
Hugging Face Blog精选AI 评分6565 Liquid AI 发布 LFM2.5-DSpark 草稿模型,推理吞吐最高提升 3.2x
Liquid AI 为 LFM2.5-1.2B-Instruct、LFM2.5-2.6B 和 LFM2.5-8B-A1B 三个模型发布 DSpark 草稿模型 checkpoint,通过投机解码在不改变输出质量的前提下加速推理,GPU 吞吐最高提升 3.18x,端侧最高 2.87x。
推荐理由:官方为 LFM2.5 三款模型发布 DSpark 草稿模型,给出从 H100 到 MacBook 的实测加速数据和开源接入方式。
@fofrAI@fofrAIAI 评分2222 过去两个月,情况真的变了。现在任何人都能做硬件。你可以给它刷入新固件,可以编写自己的操作系统。你的硬件如今真正属于你了,这在以前从未有过。去破解、去构建、去连接一切吧。
Hugging Face Blog精选AI 评分8181 Meta 发布开源多模态模型 Muse Glimmer-30B,主打本地智能体场景
Meta 发布从 Muse 蒸馏而来的 30B 多模态模型 Muse Glimmer,采用 Apache 2.0 许可,面向本地隐私场景的智能体用途。模型由 2B ViT 视觉编码器和 28B 文本解码器组成,支持图像、视频、多模态工具调用和目标检测,并附带基于 DFlash 的可选投机解码。
推荐理由:原文给出架构组成、基准对比和各推理框架的 day-0 用法,读者可以据此评估本地部署的可行路径。
Google ResearchAI 评分4545 Google 用冻结式多 Token 预测加速 Pixel 上的 Gemini Nano v3
Google Research 推出新架构,将多 Token 预测(MTP)头挂载到已冻结的 Gemini Nano v3 上,在 Pixel 9 上实现 50% 以上的推理加速。
Claude BlogAI 评分4747 Anthropic 为 Apple Foundation Models 框架推出 Claude Swift 包
Anthropic 发布 Swift 包,让 Apple 开发者通过 Foundation Models 框架调用 Claude 处理多步推理、代码生成等复杂任务,并支持联网搜索与代码执行。
Mistral AI精选AI 评分7777 Mistral 发布 Mistral 3 系列模型,含 Mistral Large 3 与 Ministral 3,均以 Apache 2.0 开源
Mistral 发布新一代模型系列 Mistral 3,包括 14B/8B/3B 的 Ministral 3 小模型和 sparse MoE 架构的 Mistral Large 3(41B 激活、675B 总参数),全部以 Apache 2.0 许可开源,提供 base、instruct 和 reasoning 变体。
推荐理由:官方公布完整模型规格、开源许可和部署路径,读者可据此评估在自建或端侧场景的可用性。