Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸
#端侧
#端侧
今日 3 条
Chubby♨️@kimmonismusAI 评分4545
引用David Stout@Davidstout
Chubby♨️@kimmonismusAI 评分4343
elvis@omarsar0AI 评分4848
引用David Stout@DavidstoutHalf a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸
Hugging Face BlogAI 评分5959 Liquid AI 发布 LFM2.5-VL-3B 的 DSpark 草稿模型,解码最高提速 3.13x
Liquid AI 为视觉语言模型 LFM2.5-VL-3B 发布实验性 DSpark 草稿模型,通过投机解码在不改变输出质量的前提下加速推理。解码最高提速 3.13x(M5 Max 设备端)和 2.66x(H100),端到端最高提升 2.62x 和 2.27x。
OpenBMB@OpenBMBAI 评分3838


inclusionAI Hugging Face modelsAI 评分5959 inclusionAI 发布 Ling-3.0-tiny:7.9B 总参数、1.3B 激活的混合推理 MoE 模型
inclusionAI 发布轻量混合推理 MoE 模型 Ling-3.0-tiny,总参数 7.9B,每 token 仅激活 1.3B,采用 3:1 KDA-MLA 混合线性架构加 128 专家稀疏 MoE FFN。
Hugging Face Blog精选AI 评分6565 Liquid AI 发布 LFM2.5-DSpark 草稿模型,推理吞吐最高提升 3.2x
Liquid AI 为 LFM2.5-1.2B-Instruct、LFM2.5-2.6B 和 LFM2.5-8B-A1B 三个模型发布 DSpark 草稿模型 checkpoint,通过投机解码在不改变输出质量的前提下加速推理,GPU 吞吐最高提升 3.18x,端侧最高 2.87x。
推荐理由:官方为 LFM2.5 三款模型发布 DSpark 草稿模型,给出从 H100 到 MacBook 的实测加速数据和开源接入方式。
Hugging Face Blog精选AI 评分8181 Meta 发布开源多模态模型 Muse Glimmer-30B,主打本地智能体场景
Meta 发布从 Muse 蒸馏而来的 30B 多模态模型 Muse Glimmer,采用 Apache 2.0 许可,面向本地隐私场景的智能体用途。模型由 2B ViT 视觉编码器和 28B 文本解码器组成,支持图像、视频、多模态工具调用和目标检测,并附带基于 DFlash 的可选投机解码。
推荐理由:原文给出架构组成、基准对比和各推理框架的 day-0 用法,读者可以据此评估本地部署的可行路径。
Google ResearchAI 评分4545 Google 用冻结式多 Token 预测加速 Pixel 上的 Gemini Nano v3
Google Research 推出新架构,将多 Token 预测(MTP)头挂载到已冻结的 Gemini Nano v3 上,在 Pixel 9 上实现 50% 以上的推理加速。
Mistral AI精选AI 评分7777 Mistral 发布 Mistral 3 系列模型,含 Mistral Large 3 与 Ministral 3,均以 Apache 2.0 开源
Mistral 发布新一代模型系列 Mistral 3,包括 14B/8B/3B 的 Ministral 3 小模型和 sparse MoE 架构的 Mistral Large 3(41B 激活、675B 总参数),全部以 Apache 2.0 许可开源,提供 base、instruct 和 reasoning 变体。
推荐理由:官方公布完整模型规格、开源许可和部署路径,读者可据此评估在自建或端侧场景的可用性。