X:Kim
@kimmonismus · X
切换来源
@kimmonismus@kimmonismusAI 评分3232 
@kimmonismus@kimmonismusAI 评分6262 @kimmonismus@kimmonismus精选AI 评分6767
引用@ArtificialAnlys@ArtificialAnlysMeta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62) Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh) ➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available ➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh) ➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh) Other model details (xhigh variant): ➤ Context window: 1M tokens, unchanged from Muse Spark 1.2 ➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Input modalities: text, image, video ➤ Availability: Meta's first-party API and Muse Code
推荐理由:作者把前沿模型竞争格局的变化讲清楚,并指出中国开源权重模型已贴近第一梯队,可与2025年的撞墙争论对照。
@kimmonismus@kimmonismusAI 评分2727 @kimmonismus@kimmonismusAI 评分1818 Fable 5 在 DeepSWE 上达到 70%。据我所知,Fable 5.1 还没有评测。https://t.co/ESjrlxywLR

@kimmonismus@kimmonismusAI 评分2727 天啊 Muse Spark 1.3 它在 DeepSWE 上以 75.4% 拿下第一名,甚至超过了 GPT-5.6 Sol 和 Fable 5。 这是怎么回事?!

@kimmonismus@kimmonismusAI 评分1313 我不想再挑起争论了,但我真希望他们明天就发布 Astra。用 Fable 5.1 我什么工作都完成不了。https://t.co/8Gz5tsvyBO

@kimmonismus@kimmonismusAI 评分66 @kimmonismus@kimmonismusAI 评分5454 
@kimmonismus@kimmonismusAI 评分3434 
@kimmonismus@kimmonismusAI 评分4141 Gemini 3.8 Flash:DeepSWE 上 73%,价格非常香!Google 这次真在发力!Gemini 4 Pro 值得期待
引用@kimmonismus@kimmonismusGemini 3.8 Flash benchmarks. And holy cow! Flash outperforms 5.6 sol and opus 5 on terminal bench 2.1, HLE and much more. Google is back! https://t.co/Vn9roSApoS
@kimmonismus@kimmonismusAI 评分66 @kimmonismus@kimmonismusAI 评分00 @kimmonismus@kimmonismusAI 评分3737 
@kimmonismus@kimmonismusAI 评分3232 
@kimmonismus@kimmonismusAI 评分55 @kimmonismus@kimmonismusAI 评分4444 
@kimmonismus@kimmonismusAI 评分2020 我删除了那条说 GPT-Astra 明天发布的推文。关于消息来源的可靠性存在相互矛盾的说法,我不想传播假新闻。这并不意味着它最终不会明天到来;只是不确定性大得多。
@kimmonismus@kimmonismusAI 评分11 @DanDr1s @synthwavedd @ChrisGPT @chetaslua @M1Astra 该给的致谢要给:https://t.co/6LcH7o6kgK
@kimmonismus@kimmonismusAI 评分55 @kimmonismus@kimmonismusAI 评分4040 引用@elonmusk@elonmuskGrok 4.7 comes out in 10 days https://t.co/ZSXmzVFqB1
@kimmonismus@kimmonismusAI 评分1010 @kimmonismus@kimmonismus精选AI 评分6666
引用@Alibaba_Qwen@Alibaba_Qwen🚀Qwen3.8-Max just got upgraded. Meet Qwen3.8-Max-0902! 2.4T parameters. 1M context tokens. Built for real world complexity. Further post trained on Coding & Cowork, Qwen3.8-Max-0902 now delivers stronger performance across complex enterprise tasks, scientific research, and long horizon workflows. 💰Pricing per 1M tokens: $2 input, $6 output. $0.17 explicit cache hit, $0.25 implicit cache hit. Now live via API on QwenCloud. Come try it! 🙌 API: https://t.co/dq3WgMk980
推荐理由:材料给出 Qwen3.8-Max-0902 的参数、上下文与定价,作者由此讨论版本迭代节奏和中美实验室差距的缩小。
@kimmonismus@kimmonismusAI 评分5353 


@kimmonismus@kimmonismusAI 评分4141 
@kimmonismus@kimmonismusAI 评分3232 引用@kimmonismus@kimmonismusLiterally unusable. The rate limits are absurd. Oh, and by the way, Fable’s automatic continuation is bugged and doesn’t even work. I honestly don’t know why I still bother using Claude at this point. 5.6 is simply better overall anyway. Give me GPT-Astra and im fine. its so frustrating. seriously. oh, and btw. For subscription users, Anthropic has not announced lower prices or higher usage limits regarding Fable 5.1s efficency gains; the savings explicitly apply “wherever usage is billed by token,” so greater efficiency within Pro or Max subscriptions possible not gonna happen.
@kimmonismus@kimmonismusAI 评分1818 
@kimmonismus@kimmonismusAI 评分4949
引用@kimmonismus@kimmonismusOpenAI’s unreleased Astra model found two V8 zero-days during testing, and used them in an exploit chain with little human help. In their new blogpost, OpenAI wrote that in separate expert assessments, Astra compromised a hardened browser, escaped its sandbox and executed commands on the host. It also chained several operating-system vulnerabilities to move from an unprivileged account to root. OpenAI has classified Astra as “Critical” for cybersecurity, the first of its models to reach that threshold. OpenAI paused parts of Astra’s training after the Hugging Face incident, but restarted the main frontier RL run on August 28 under stricter controls.
@kimmonismus@kimmonismusAI 评分55 @kimmonismus@kimmonismus精选AI 评分8080 

推荐理由:材料给出 Astra 在漏洞挖掘与提权测试中的具体结果和风险定级,可了解前沿模型在网络安全场景的能力边界。
@kimmonismus@kimmonismusAI 评分4040 

@kimmonismus@kimmonismus精选AI 评分6565 


推荐理由:文中并列给出 Fable 5.1 的定价、缓存成本与多项智能体基准变化,便于对比升级前后的取舍。
@kimmonismus@kimmonismusAI 评分2929 
@kimmonismus@kimmonismusAI 评分2323 Fable 5.1 上线了,连我在德国都能用。冲!https://t.co/abtJbeADcA

@kimmonismus@kimmonismusAI 评分2323 来了,也许 Fable 5.1 今天终究会官宣。 已加入 Claude Code 2.1.257 https://t.co/OOJmZiGUwe https://t.co/rYf9Hyzc81

@kimmonismus@kimmonismusAI 评分2525 Fable 5.1 已在支持文档中被正式提及。发布在即。https://t.co/asZdLMzFxd https://t.co/TTZw4psN3r

@kimmonismus@kimmonismusAI 评分1414
@kimmonismus@kimmonismusAI 评分66 @kimmonismus@kimmonismusAI 评分77 @kimmonismus@kimmonismusAI 评分4343 


