跳到正文

#开源生态

今日 12 条
今天10月2日周五
  1. Hugging Face Daily Papers47

    Omni-Embed-Mini:通过密集蒸馏绑定多模态而不遗忘文本

    Omni-Embed-Mini 以 0.9B 参数将文本、语音、音频、图像、视频和富文本文档映射到统一余弦空间,且不更新任何文本侧参数。其核心思路是无需独立嵌入模型作为教师,直接以冻结骨干网络对密集级联标题的嵌入作为目标,配合 Matryoshka SigLIP 对比损失与在线混合难负例挖掘器完成对齐。

  2. Chubby♨️45

    webAI 发布 3.66B 参数形式逻辑模型 TwIL-LM3-Pro,可在笔记本本地运行。其综合逻辑评测与 Qwen3-8B 持平,参数量不足后者一半,并在全部六项形式逻辑任务上领先 VibeThinker-3B。该模型基于 IBM Granite 4.2 后训练,Q4 GGUF 权重仅 2.09 GiB,可通过 llama.cpp 本地推理。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  3. Hacker News popular via buzzing.cc76

    DeepSeek Harness 开启全球公开预览并开源

    DeepSeek Harness 进入全球公开预览并开源,基于 Cordis 的“一切皆插件”架构,可作为桌面应用运行或从代码启动 Web UI。它支持日常办公、编码、研究、后台任务,可通过“Creator mode”在聊天中创建插件,用 npx @deepseek-ai/dsh web 一条命令启动,源码在 github.com/deepseek-ai/deepseek-harness。

    推荐理由:原文给出 DeepSeek Harness 的插件架构、安装方式和适用场景,读者可据此评估是否纳入自己的工作流。

  4. elvis48

    webAI 开源 3.6B 参数模型 TwIL-LM3-Pro,可在普通电脑本地运行,BIG-Bench Hard 得分 95.4,远超 Qwen3-8B 的 63.7。其训练配方为:形式逻辑微调后将权重合并回基座模型,再用程序化验证器做 RL,逻辑分数提升且通用推理保持稳定。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  5. TechCrunch · AI56

    AWS 发布开源决策模型 Strands Decider 2B,基于 Qwen3.5-2B

    AWS 发布开源决策模型 Strands Decider 2B,灵感来自 TypeSafe 的 Jev,可在预设选项间高速低成本地做选择并给出置信度。模型完全开源、可本地运行,由 Amazon 杰出工程师 Marc Brooker 的内部项目改进而来,基于 Qwen3.5-2B 的架构但不生成文本,而是输出校准后的选择,同一周 OpenAI 也宣布了类似产品。

10月1日周四
  1. ViggleAI44

    ✨ Viggle Turbo v0.3 来了! 全新的 9-step 模式带来更干净的画面、更精细的细节,以及更好的小文字渲染。 现在在 ComfyUI 中更易使用——提供 LoRA 或单文件 int8/fp8/GGUF 模型。

    引用Yun Chen@t_mux

    Viggle Turbo v0.3 for Qwen-Image-2.1 is out! - Less grain than v0.2.1, a touch softer - New 9-step mode: finer detail, small text - ComfyUI: LoRA or single-file int8/fp8/GGUF Model: https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

  2. Aravind Srinivas53

    Perplexity 开源其上下文嵌入模型,该模型在 turbopuffer 的 context-bench 上表现最佳。引用内容显示 pplx-embed-v2-context-9b-preview 采用整篇文档视野下编码每个文本块的新训练方式,在 ConTEB 和 turbopuffer 的 context-bench 上创下 SOTA,详见 https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage

    引用Perplexity@perplexity_ai

    We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage

  3. Demis Hassabis66

    Google DeepMind 推出蛋白质水印方法 SynthID Bio,成功合成既有功能又带水印的 AI 设计蛋白质,成果发表于 Nature。作者称生物安全是 AI 时代最紧迫挑战之一,并开源 SynthID Bio 工具供研究社区使用。

    引用Pushmeet Kohli@pushmeet

    Very happy to announce that our team @GoogleDeepmind has pushed the boundaries of generative biology, achieving the successful synthesis of AI-designed proteins that are both functional and watermarked. This proof-of-concept watermarking of the building blocks of life is enabled by SynthID Bio, our new protein watermarking method. It is designed to safeguard the new era of AI-powered generative biology and strengthen global biosecurity. You can read my thoughts here on why watermarking AI-designed proteins is an important research breakthrough: https://x.com/pushmeet/status/2105314763148321102

    推荐理由:AI 设计蛋白质首次实现功能与水印兼具并发表于 Nature,同时开源工具,读者可关注生物安全水印路线的实际落地。

  4. Thomas Wolf34

    ESM-2 于 2022 年发布。 它至今每月仍有数十万次下载。 这能持续,全靠有人在维护它底层的软件。 @huggingface 🤝 @os4science 正联手找出这些库,并支持它们背后的维护者 🧬 https://os4science.org/news/hugging-face-open-source-for-science-fund/

    引用Open Source for Science Fund@os4science

    We're joining forces with @huggingface to identify the software libraries that scientific model contributors rely on most and explore opportunities to support the maintainers behind them. https://os4science.org/news/hugging-face-open-source-for-science-fund/

  5. Anthropic Research73

    物理学家 Schwartz 分享用 Claude 与 BootLoops 做跨领域定量科学的方法与成果

    物理学家 Matthew Schwartz 发文介绍一种 AI 加速科研的新思路:不再对抗模型短板,而是寻找适合当前 LLM 的 "Claude-shaped" 问题,并开源了定量科学计算工具包 BootLoops。

    推荐理由:作者复盘了寻找 AI 擅长问题并联合领域专家雕琢结果的完整方法,这套协作模式对科研用户可直接借鉴。

9月30日周三
  1. clem 🤗79

    Hugging Face CEO Clément Delangue 发文称在被 NVIDIA 收购后收到了数千条招聘私信,@bot 无法自动分析私信,需几天时间逐一查看。他表示未获回复不代表被否定,目前只聚焦特别匹配的人选,建议同时到 https://apply.workable.com/huggingface 申请具体职位,并感谢大家继续推动开源 AI。

    引用clem 🤗@ClementDelangue

    getting acquired by @nvidia = hugging face can now hire people we couldn't as a small startup and give them a decade to make open-source AI win! if you're one of them, my dms are open

    推荐理由:作者回应了被 NVIDIA 收购后的招聘进展,说明了私信道申请的处理方式和官网投递渠道,对有意加入者有直接参考。

  2. Tianyi Cui38

    今天推荐一个功能很丰富的 DSH 上下文管理插件 dsh-context: https://github.com/bowenliang123/dsh-context

    引用Tianyi Cui@tianyi

    从 DeepSeek 官方 API 处统计的数据来看,约有 60% 的 DeepSeek Harness 用户使用了至少一个第三方插件。第三方插件是 DeepSeek Harness 用户体验中最具特色且不可缺少的一部分。DeepSeek Harness 团队将持续支持第三方插件生态的繁荣发展,并推动插件 API 趋于稳定,在将来减少和尽量避免破坏性更新。 接下来的几天我个人将每天推荐一个优质的 DSH 第三方插件,欢迎 DSH 插件作者在本 thread 下自荐。我会结合插件质量及后台实际统计到的插件使用量择优推荐。 DeepSeek Harness 团队祝大家中秋快乐阖家幸福! (注:在用户使用官方 API 及模型时,DSH 会向官方 API 上报实际使用的插件包名和版本。此类上报不额外消耗 tokens。)

9月29日周二
  1. Thomas Wolf56

    modded-nanogpt 传入新的历史纪录 39.9 秒,较此前 67.6 秒快 27.7 秒,核心思路是在单个 flop 级别做稀疏优化而非只优化矩阵乘法。主要手段包括采样 softmax(约 8 秒)、稀疏 n-gram 嵌入更新与优化器状态、稀疏通信、最后 300 步 EMA(约 4 秒)、新优化器 Anvil2(约 1 秒)等,稀疏嵌入参数扩展到 65B,占本次提升的 25%。详见 https://github.com/KellerJordan/modded-nanogpt/pull/360 和 https://hyperstition.cc/training-nanogpt-in-39-9-seconds。

    引用Larry Dial@classiclarryd

    New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds

  2. clem 🤗77

    AMD 宣布欢迎 World Labs 和李飞飞加入 AMD,双方计划结合 World Labs 在 AI 与世界模型方面的专长与 AMD 的算力能力,推动 AI 未来并强化开放 AI 生态。Hugging Face CEO Clément Delangue 转发该消息并祝贺,期待双方未来数年的成果。

    引用Lisa Su@LisaSu

    So excited to welcome @theworldlabs and @drfeifei to the @AMD family! I’ve always been a huge fan of Fei-Fei and her pioneering research in AI. Together, we’ll combine World Labs’ deep expertise in AI and world models with AMD’s compute leadership to power the future of AI and strengthen the open AI ecosystem. Can’t wait for all we’ll accomplish!

    推荐理由:AMD 收购 World Labs 与李飞飞的消息结合 World Labs 专注世界模型与开放生态的定位,读者可了解这次结合对 AI 开源生态的影响。

  3. Google Cloud: Databases65

    Google Cloud 分析创业公司为何需要在前沿 API 之外搭配 Gemma 4 开源模型

    Google Cloud 发文主张创业公司采用“复合 AI 栈”,用开源的 Gemma 4 处理边缘执行、高吞吐分流、任务微调和垂直场景,把 Gemini 留给复杂推理。

    推荐理由:文章用三个创业案例和四类工作负载说明开源模型与前沿 API 搭配的架构取舍,适合正在做模型选型的团队参考。

9月28日周一