This @Bloomberg story is BS. Argon has been my daily driver for a while and it's been a great experience. It's particularly awesome at agentic debugging besides day-to-day coding tasks.
#模型发布
#模型发布
今日 1 条
Chubby♨️@kimmonismusAI 评分2727引用Aditya Timmaraju@tadityasrinivas
赵纯想@chunxiangaiAI 评分2626好了,发布前灰度赏赐我,发布后不给用了。史上体感最明显的降智,从神降到 3.8 flash,跳楼机?蹦极?
引用赵纯想@chunxiangai重大信息差: 赶紧把 Google 反重力下回来,沐浴焚香,洗洗干净,迎接神的降临。
Dongxi 东锡 NLP@dongxi_nlpAI 评分3131
Deedy@deedydasAI 评分3131我们正处在晚期 AI 模型资本主义阶段,你知道前沿模型必须在一堆基准测试上获胜才能发布。这些数字毫无意义。 值得信任的是价格。 如果定价高,那就是好模型。如果不高,那就是刷榜的。
Yuchen Jin@Yuchenj_UWAI 评分2222Fast:快 2 倍,价格 2 倍。 Ultrafast:快 8 倍,价格 6 倍。 也许我们应该推出一些 Ultra-ultrafast 开源模型端点?

jietang@jietangAI 评分2424你确定吗?找到最优模型规模很棘手:数据量、激活参数量、环境数量,以及目标推理成本。模型性能还取决于许多其他因素,每个因素都会带来自身的变数。
引用Charlie O'Neill@oneill_cFable is probably ~2-2.5T parameters, not 10T. Kimi K3 is 2.8T params, trained on maybe 20–30k Blackwell-equivalents. It lands within spitting distance of Fable 5 in terms of capabilities (5, not 5.1). Anthropic has far more compute than Moonshot, better rl environments, better architecture and better optimizers and all of that adds to capability per parameter. So if Fable is only slightly ahead of K3 with this in mind, it's almost certainly a smaller model. GPT-5.5 and 5.6 are smaller still (I'll say more on that later)
Every latest articlesAI 评分4444 实测专为修复 AI 写作而生的 AI 模型:文风更难预测,但未必更好
一家新实验室认为 AI 文风千篇一律是训练问题,并据此推出了一款专注写作的模型。实测显示,它生成的文字更难预测,但质量未必更好。作者称主要实验室在写作上的模型进展已停滞,而 OpenAI、Anthropic、Google 或更侧重编程方向。
Nathan Lambert: Interconnects精选AI 评分7171 Nathan Lambert 解析 GLM-5.3 为何能紧跟前沿
Z.ai 发布 GLM-5.3,目前仅在 coding plan 提供,两周内将开放权重到 Hugging Face。作者认为其与 GLM-5.2 同底座、靠大幅扩展后训练提升成绩,在部分智能体编码基准上超越 Kimi K3 甚至个别超越 Claude Fable 5 或 GPT-5.6-Sol,参数约 750B。
推荐理由:作者给出了对 GLM-5.3 成绩来源的解释框架,包括发布节奏、后训练策略和 RL 数据产业等背景,可用于理解中美前沿模型竞争的成因。