跳到正文

#模型发布

今日 1 条
今天10月2日周五
  1. Chubby♨️27

    今日最佳消息:@GoogleDeepMind 的一位 Senior Staff Research Engineer 称 Bloomberg 的报道是“胡说八道”。 Bloomberg 声称,虽然 Gemini 4 “在广泛用于衡量模型效能的基准测试中表现良好,但当员工真正将其投入工作时,表现就不那么好了。” 所以是的,这让我们更有理由期待一次出色的发布。

    引用Aditya Timmaraju@tadityasrinivas

    This @Bloomberg story is BS. Argon has been my daily driver for a while and it's been a great experience. It's particularly awesome at agentic debugging besides day-to-day coding tasks.

10月1日周四
9月30日周三
9月10日周四
  1. jietang24

    你确定吗?找到最优模型规模很棘手:数据量、激活参数量、环境数量,以及目标推理成本。模型性能还取决于许多其他因素,每个因素都会带来自身的变数。

    引用Charlie O'Neill@oneill_c

    Fable is probably ~2-2.5T parameters, not 10T. Kimi K3 is 2.8T params, trained on maybe 20–30k Blackwell-equivalents. It lands within spitting distance of Fable 5 in terms of capabilities (5, not 5.1). Anthropic has far more compute than Moonshot, better rl environments, better architecture and better optimizers and all of that adds to capability per parameter. So if Fable is only slightly ahead of K3 with this in mind, it's almost certainly a smaller model. GPT-5.5 and 5.6 are smaller still (I'll say more on that later)

8月24日周一
8月15日周六
  1. Nathan Lambert: Interconnects71

    Nathan Lambert 解析 GLM-5.3 为何能紧跟前沿

    Z.ai 发布 GLM-5.3,目前仅在 coding plan 提供,两周内将开放权重到 Hugging Face。作者认为其与 GLM-5.2 同底座、靠大幅扩展后训练提升成绩,在部分智能体编码基准上超越 Kimi K3 甚至个别超越 Claude Fable 5 或 GPT-5.6-Sol,参数约 750B。

    推荐理由:作者给出了对 GLM-5.3 成绩来源的解释框架,包括发布节奏、后训练策略和 RL 数据产业等背景,可用于理解中美前沿模型竞争的成因。