跳到正文

#推理

今日 5 条
今天10月2日周五
  1. Chubby♨️45

    webAI 发布 3.66B 参数形式逻辑模型 TwIL-LM3-Pro,可在笔记本本地运行。其综合逻辑评测与 Qwen3-8B 持平,参数量不足后者一半,并在全部六项形式逻辑任务上领先 VibeThinker-3B。该模型基于 IBM Granite 4.2 后训练,Q4 GGUF 权重仅 2.09 GiB,可通过 llama.cpp 本地推理。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  2. TechCrunch · AI63

    Google 发布 Gemini 4 Argon,称其为迄今最强模型

    Google(Alphabet)发布新模型 Gemini 4 Argon,主打防御性网络安全,称其可自主发现、验证并修复关键软件漏洞,目前仅通过 Fairwind 安全计划向部分网络安全合作伙伴开放。该模型也用于编码、调试和代码库迁移等日常工程工作,并称在多项基准上显著领先 GPT-6 Astra 与 Anthropic 的 Fable 和 Opus。

  3. Chubby♨️42

    来了:大量用户现在报告他们的查询正被路由到 Fable 5.5! - 不带网页搜索的查询现在能返回最新结果。 - 初步 SVG 测试远优于 Fable 5.1。 这只是时间问题。准备好,世界上最好的模型即将发布。

    引用Chetaslua@chetaslua

    🚨 Fable 5.5 is auto routing on web , this is the screenshot it edited for x without even prompted he knows tibo check https://claude.ai see if you are getting routed or not

  4. elvis48

    webAI 开源 3.6B 参数模型 TwIL-LM3-Pro,可在普通电脑本地运行,BIG-Bench Hard 得分 95.4,远超 Qwen3-8B 的 63.7。其训练配方为:形式逻辑微调后将权重合并回基座模型,再用程序化验证器做 RL,逻辑分数提升且通用推理保持稳定。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

10月1日周四
  1. The Decoder79

    Google 发布 Gemini 4 Argon,追赶 OpenAI 与 Anthropic 但未取得明确领先

    Google 发布新旗舰模型 Gemini 4 Argon,是其七个多月来首款前沿模型,Artificial Analysis 测试中得 53 分,与 GPT-6 Astra (max)、Claude Fable 5.1 持平,但仍落后 Claude Opus 5.5 的 58 分。

    推荐理由:原文汇总了第三方测试与定价细节,指出 Gemini 4 Argon 缩小差距但未领先,且单价优势来自低 token 价格而非效率。

  2. Karina52

    Google DeepMind 发布新前沿模型 Gemini 4 Argon,面向编码、企业知识工作和网络安全防御等复杂工作流,即日起通过 Fairwind Program 向部分受信任测试者开放。作者引述其 PostTrainBench 得分 45.3%,高于 Gemini 3.1 Pro 的 21.99% 和 GPT-6 Astra 的 44.3%;评测表还显示其在自动化与智能体编码等多项基准领先,但在 FrontierSWE v2、Terminal-Bench 4.0 等项落后于对比模型。

    引用Google DeepMind@GoogleDeepMind

    Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.

  3. fofr33

    非常激动地分享,Gemini 4 Argon 即将到来。迫不及待想尽快跟大家分享更多内容。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  4. The Verge · AI73

    Google 发布 Gemini 4 Argon,初期仅限受信任的网络防御者使用

    Google 发布新一代前沿模型 Gemini 4 Argon,称其在软件工程、法律金融等企业知识工作和网络安全防御方面具有前沿性能。初期仅向一组受信任的网络防御者开放,Google 正参与美国政府预发布模型访问的自愿流程并逐步扩大访问。模型已用于 Google 内部工作流,如大规模代码库迁移;Google 将在更广泛发布前加强防范滥用和提示词注入攻击、监测错位等安全措施。

9月29日周二
9月26日周六
  1. Anthropic65

    Anthropic 发文称,Claude 在收到单个九圈问题提示词后,在 Claude Science 中基本无人监督地运行数天,用 Dixon 等人的方法完成求解,总成本几千美元,突破了此前八圈的纪录(平面 N=4 超杨-米尔斯简化模型)。物理学家 Lance Dixon 独立验证了结果,von Hippel 为该博客撰写了经历回顾。

    推荐理由:九圈散射振幅计算由 Dixon 独立验证,计算成本仅几千美元,为学界评估 AI 科研能力提供了一个可核验的案例。

9月24日周四
  1. inclusionAI Hugging Face models60

    inclusionAI 开源 Ling-mini-2.0:16B 总参数 MoE 模型,激活仅 1.4B

    inclusionAI 开源 Ling 2.0 系列首个模型 Ling-mini-2.0,总参数 16.26B、每 token 激活 1.4B(非嵌入 789M),采用 1/32 激活比 MoE 架构,称可达到 7–8B dense 模型的等效性能,在 H20 上简单 QA 场景生成速度超过 300 token/s,支持 128K 上下文(YaRN)。

    推荐理由:官方发布了完整的参数配置、推理速度、训练吞吐和预训练检查点,读者可以据此评估小激活 MoE 在端侧和继续训练中的可用性。

  2. inclusionAI Hugging Face models62

    inclusionAI 开源 Ling-flash-2.0:100B 总参数、6.1B 激活的 MoE 模型

    inclusionAI 正式开源 Ling 2.0 架构下的第三个 MoE 大语言模型 Ling-flash-2.0,总参数 100B、激活参数 6.1B(非嵌入 4.8B),基于 20T+ tokens 数据训练并经 SFT 和多阶段强化学习。

    推荐理由:官方给出参数结构、基准对比和推理速度数据,读者可据此评估小激活 MoE 替代 40B 稠密模型的可行性。

  3. inclusionAI Hugging Face models68

    inclusionAI 开源发布万亿参数思考模型 Ring-1T

    inclusionAI 正式发布开源思考模型 Ring-1T,总参数 1 万亿、激活 50B,基于 Ling 2.0 架构和 Ling-1T-base,上下文经 YaRN 扩展至 128K,权重可在 Hugging Face 与 ModelScope 下载,并支持 Ling Chat 和 ZenMux 体验与 API 调用。

    推荐理由:官方发布开源万亿参数思考模型,给出 IMO 与 ICPC 实测结果和 Icepop、ASystem 训练细节,便于评估其推理与部署价值。

  4. inclusionAI Hugging Face models62

    inclusionAI 开源万亿参数思考模型 Ring-2.5-1T

    inclusionAI 发布开源万亿参数思考模型 Ring-2.5-1T,采用 1:7 MLA + Lightning Linear Attention 混合线性注意力架构,激活参数从 51B 增至 63B。

    推荐理由:原文给出了混合线性注意力的具体效率数字和多项基准成绩,读者可以据此评估它在长程智能体场景中的实际取舍。

  5. inclusionAI Hugging Face models62

    inclusionAI 开源 Ling-2.5-1T:1T 总参数、63B 激活的即时模型,支持 1M 上下文

    inclusionAI 发布并开源 Ling-2.5-1T,总参数 1T、激活 63B,预训练语料扩至 29T tokens,经 YaRN 外推支持最长 1M tokens 上下文。

    推荐理由:官方模型卡给出架构、token 效率和长上下文评测细节,可帮助读者评估这款万亿参数即时模型对现有部署工作流的适配。

  6. inclusionAI Hugging Face models60

    inclusionAI 发布万亿参数推理模型 Ring-2.6-1T

    inclusionAI 发布万亿参数旗舰推理模型 Ring-2.6-1T,主打真实生产环境中的 Agent 执行与复杂推理,上下文长度由 128K 扩展到 256K(YaRN),采用 MIT License 开源。

    推荐理由:官方发布页给出 Agent 执行、推理档位和异步 RL 训练的具体做法与基准数字,读者可据此评估其在生产场景的适配价值。

9月23日周三
  1. Greg Brockman79

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna 两款更快、更实惠的模型,基于 GPT-6 Astra 的技术积累。两款模型在专业工作、事实性、编码、computer use 和对齐方面延续 Astra 的 SOTA 表现,同时缓存和推理效率提升使 API 价格比 GPT-5.6 促销价低 50%。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:原文由当事方宣布两款新模型及降价幅度,读者可以据此了解 GPT-6 系列的能力分工与成本变化。

  2. Noam Brown82

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna,性能优于 GPT-5.6 且 API 价格低 50%。Luna 现为 $0.10 输入 / $0.50 输出每 1M tokens,这是继 7 月底 Luna 降价 80% 之后的又一次下调,两个月内输出价格从 $6 降至 $0.50。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:作者以当事方身份给出模型、降价幅度和具体价格,可据此比较 GPT-6 系列的成本变化。

9月22日周二
  1. Andrew Milich55

    Grok 4.7 发布,官方称在同等价格和速度下较 Grok 4.6 有明显提升。作者推荐在 Grok Build 和 Cursor 中以高 TPS 尝试,称其在编码、工程工作和 3D 方面表现出色。附表显示 Grok 4.7 xHigh 输入 $2/百万 token、输出 $6/百万 token,与 Grok 4.6 相同;Cursor Bench 4.0 得分 46.3%(4.6 为 40.4%),EEBench 64.0%(53.0%),Harvey Legal Agent 19.6%。

    引用SpaceXAI@SpaceXAI

    Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.

  2. Anthropic Newsroom90

    Anthropic 发布 Claude Opus 5.5,成本较 Opus 5 降低 40%

    Anthropic 发布 Claude Opus 5.5,为 Claude 5.5 家族首款模型,官方称其表现与 Claude Fable 5.1 相当,运行成本较 Opus 5 降低 40%,输入和输出 token 价格为 $4 和 $20 每百万,缓存读取 $0.20 每百万(降低 60%),输出速度快 30% 以上。

    推荐理由:官方给出完整基准、价格与安全评估细节,读者可据此比较 Opus 5.5 在成本与智能体编码上的实际变化。

9月19日周六
9月14日周一