Meta 推出新的个人 AI 助手 Muse,官方称其常驻运行、速度很快,能使用浏览器并连接用户的应用,设计上注重安全。该助手现已开放试用,试用入口为 https://t.co/n7swQh9v6C。
X:Alexandr Wang(Scale AI 创始人/Meta 首席 AI 官)
@alexandr_wang · X
切换来源
@alexandr_wang@alexandr_wangAI 评分5858 
@alexandr_wang@alexandr_wangAI 评分1515
@alexandr_wang@alexandr_wangAI 评分1111 @alexandr_wang@alexandr_wangAI 评分1010 @alexandr_wang@alexandr_wangAI 评分1414 muse spark 1.3 max 和 gpt-6 astra 的对比不错 https://t.co/AyC99H6agD
@alexandr_wang@alexandr_wangAI 评分3333 Vals AI 的 Muse Spark 1.3 Max,与 Claude Fable 5 和 GPT-5.6 Sol 竞争,但便宜 4-8 倍 https://t.co/hPEk5hKv0R
@alexandr_wang@alexandr_wangAI 评分22 @alexandr_wang@alexandr_wangAI 评分4444 引用@ArtificialAnlys@ArtificialAnlysAnnouncing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index v4.2 changelog: + AA-Briefcase, our agentic knowledge work evaluation with a private test set + @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! Intelligence Index v4.2 changes in detail: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5. ➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure. Key results: ➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google ➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier ➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
@alexandr_wang@alexandr_wangAI 评分44 @alexandr_wang@alexandr_wangAI 评分1313 4/ 或者直接在 muse code 里试用。把这段复制粘贴到你的终端: curl -fsSL https://t.co/TFXsXfJKW8 | bash
@alexandr_wang@alexandr_wangAI 评分55 @alexandr_wang@alexandr_wangAI 评分99 @alexandr_wang@alexandr_wangAI 评分1111 2/ 提醒一下,我们是在完成安全测试之后才发布这个的。 期待大家试用!https://t.co/jsFNaqWBIL

@alexandr_wang@alexandr_wangAI 评分3838 



@alexandr_wang@alexandr_wangAI 评分00 @alexandr_wang@alexandr_wangAI 评分66 @alexandr_wang@alexandr_wangAI 评分1515 @alexandr_wang@alexandr_wangAI 评分22 @alexandr_wang@alexandr_wangAI 评分55
@alexandr_wang@alexandr_wangAI 评分11 @alexandr_wang@alexandr_wangAI 评分33 @alexandr_wang@alexandr_wangAI 评分2020 @alexandr_wang@alexandr_wangAI 评分66 @alexandr_wang@alexandr_wangAI 评分55 @alexandr_wang@alexandr_wangAI 评分44 @alexandr_wang@alexandr_wangAI 评分44 @alexandr_wang@alexandr_wangAI 评分22 @alexandr_wang@alexandr_wangAI 评分77 @alexandr_wang@alexandr_wangAI 评分33 @alexandr_wang@alexandr_wangAI 评分77 @alexandr_wang@alexandr_wangAI 评分5454 引用@ArtificialAnlys@ArtificialAnlysMeta's Muse Spark 1.3 (max), which is in limited preview for Meta's partners, scores 68 on the Artificial Analysis Coding Agent Index in the Muse Code harness, #2 behind only Claude Opus 5 (xhigh) in Claude Code. The variant available now, Muse Spark 1.3 (xhigh), scores 64 and costs the least per task of any agent above a 60 index score Muse Spark 1.3 (xhigh) enters the Artificial Analysis Coding Agent Index at 64 in Muse Code, up 2 points from Muse Spark 1.2 (62, August). It enters level with Grok 4.5 (high) in Grok Build (64) and behind GPT-5.6 Sol (max) in Codex (65). At $1.72 per task, it costs the least of any agent above a 60 index score, around a fifth of the cost of Claude Opus 5 (xhigh) in Claude Code ($8.17) Muse Spark 1.3 (max), which is in a limited preview stage, lands at 68 in Muse Code. It enters behind only Claude Opus 5 (xhigh) in Claude Code (68), and ahead of Claude Fable 5 (max) in Claude Code (67) and GPT-5.6 Sol (max) in Codex (65). Muse Spark 1.3 (max) is excluded from cost comparisons as Meta has not announced pricing for the limited release Claude Fable 5.1 results are in progress and will be added when complete. Congratulations @AIatMeta, @finkd, and @alexandr_wang on this result!
@alexandr_wang@alexandr_wangAI 评分77 @alexandr_wang@alexandr_wangAI 评分55 @alexandr_wang@alexandr_wangAI 评分2222 @alexandr_wang@alexandr_wangAI 评分1313 @alexandr_wang@alexandr_wangAI 评分1717