跳到正文

#端侧

今日 8 条
今天10月2日周五
  1. Chubby♨️45

    webAI 发布 3.66B 参数形式逻辑模型 TwIL-LM3-Pro,可在笔记本本地运行。其综合逻辑评测与 Qwen3-8B 持平,参数量不足后者一半,并在全部六项形式逻辑任务上领先 VibeThinker-3B。该模型基于 IBM Granite 4.2 后训练,Q4 GGUF 权重仅 2.09 GiB,可通过 llama.cpp 本地推理。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  2. TechCrunch · AI48

    Legato 推出 AI 助听眼镜 Legato Frames,起售价 999 美元

    听力科技初创公司 Legato 周四宣布,其 AI 助听眼镜 Legato Frames 正式开售,起售价 999 美元,面向轻至中度听力损失成年人。该眼镜将专利助听技术集成于镜腿,AI 系统可区分人声与背景噪音,双扬声器系统在耳旁数英寸处降低 99% 漏音,重 34 克,续航 10-12 小时,支持处方镜片和蓝牙串流。

  3. elvis48

    webAI 开源 3.6B 参数模型 TwIL-LM3-Pro,可在普通电脑本地运行,BIG-Bench Hard 得分 95.4,远超 Qwen3-8B 的 63.7。其训练配方为:形式逻辑微调后将权重合并回基座模型,再用程序化验证器做 RL,逻辑分数提升且通用推理保持稳定。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

9月30日周三
9月29日周二
  1. Google Cloud: Databases65

    Google Cloud 分析创业公司为何需要在前沿 API 之外搭配 Gemma 4 开源模型

    Google Cloud 发文主张创业公司采用“复合 AI 栈”,用开源的 Gemma 4 处理边缘执行、高吞吐分流、任务微调和垂直场景,把 Gemini 留给复杂推理。

    推荐理由:文章用三个创业案例和四类工作负载说明开源模型与前沿 API 搭配的架构取舍,适合正在做模型选型的团队参考。

9月28日周一
9月25日周五
9月24日周四
  1. Apple Machine Learning Research39

    Apple 用潜空间蒸馏压缩流式神经音频编码器,2.8 倍压缩下 WER 仅升 1.9%

    Apple 研究团队提出潜空间蒸馏方法压缩流式神经音频编码器(tokenizer),以教师模型量化前的逐帧潜表示为监督目标,仅训练学生编码器做回归,并用单层仿射层吸收师生宽度差异。在 2.8 倍压缩下,六组师生配对中有五组无需微调即保持在教师 1.9% 相对 WER 以内,并比同容量独立训练编码器相对提升 3.9%。该方法同时适用于独立预训练和与语言模型联合训练的 tokenizer。

9月23日周三
  1. elsewhere articles33

    弋途科技完成近亿元 Pre-B 轮融资,AICAR 正式上线

    弋途科技(EXTURING)近日完成近亿元 Pre-B 轮融资,由上海半导体装备材料产业投资基金与 Sands Talk Capital 联合领投,资金将用于加速 AICAR 规模化推广及自研端侧模型落地。公司已服务超 70% 国内头部车企及主流合资品牌,数十款车型量产上车,累计交付达百万级,AICAR 近期正式上线。

  2. Google Developers Blog65

    Antigravity SDK 支持本地模型,首发接入 Gemma 4 26B A4B 与 LiteRT

    Google 宣布 Antigravity SDK 支持本地工作流,首发支持通过 Google AI Edge 的 LiteRT 运行 Gemma 4 26B A4B,可完全离线提供智能体能力,建议机器配备 24GB 以上 VRAM 或统一内存。

    推荐理由:原文给出本地运行智能体的具体配置方法和混合编排演示数据,开发者可以直接照此把智能体工作流搬到本地 GPU 上。

9月22日周二
  1. Hugging Face Blog48

    oMLX 创作者 Jun Kim 加入 Hugging Face,支持 MLX 社区

    oMLX 创作者兼维护者 Jun Kim 加入 Hugging Face,全职投入 MLX 社区建设。oMLX 将从副业转为有资金支持的正式项目,继续以 Apache 2.0 开源并由 Jun 领导,目标是加快开发并更好引导贡献者。Hugging Face 计划让 oMLX 成为新想法的试验场,并推动 transformers 模型定义快速转为可被不同引擎使用的 MLX 参考实现。

9月21日周一
8月21日周五
  1. Hugging Face Blog65

    Liquid AI 发布 LFM2.5-DSpark 草稿模型,推理吞吐最高提升 3.2x

    Liquid AI 为 LFM2.5-1.2B-Instruct、LFM2.5-2.6B 和 LFM2.5-8B-A1B 三个模型发布 DSpark 草稿模型 checkpoint,通过投机解码在不改变输出质量的前提下加速推理,GPU 吞吐最高提升 3.18x,端侧最高 2.87x。

    推荐理由:官方为 LFM2.5 三款模型发布 DSpark 草稿模型,给出从 H100 到 MacBook 的实测加速数据和开源接入方式。

8月18日周二
8月10日周一
  1. Hugging Face Blog81

    Meta 发布开源多模态模型 Muse Glimmer-30B,主打本地智能体场景

    Meta 发布从 Muse 蒸馏而来的 30B 多模态模型 Muse Glimmer,采用 Apache 2.0 许可,面向本地隐私场景的智能体用途。模型由 2B ViT 视觉编码器和 28B 文本解码器组成,支持图像、视频、多模态工具调用和目标检测,并附带基于 DFlash 的可选投机解码。

    推荐理由:原文给出架构组成、基准对比和各推理框架的 day-0 用法,读者可以据此评估本地部署的可行路径。

6月27日周六
6月8日周一
12月3日周三
  1. Mistral AI77

    Mistral 发布 Mistral 3 系列模型,含 Mistral Large 3 与 Ministral 3,均以 Apache 2.0 开源

    Mistral 发布新一代模型系列 Mistral 3,包括 14B/8B/3B 的 Ministral 3 小模型和 sparse MoE 架构的 Mistral Large 3(41B 激活、675B 总参数),全部以 Apache 2.0 许可开源,提供 base、instruct 和 reasoning 变体。

    推荐理由:官方公布完整模型规格、开源许可和部署路径,读者可据此评估在自建或端侧场景的可用性。