跳到正文

Meta / Llama

Meta 的 AI 动态:Llama 开源模型系列、超级智能实验室与元宇宙之外的 AI 豪赌。

当前仅显示精选新闻
44条精选相关主题xAI / Grok开源生态行业动态

最新精选

第 21–40 条 · 共 44 条
9月3日周四
  1. Tomer Tunguz74

    Meta 新定价如何把广告换数据的模式带进 AI

    Meta 发布开源模型 Muse Spark 并推出双轨定价:标准档 muse-spark-1.3 按每百万 token 输入 $1.25、输出 $4.25 计费且承诺零数据保留,contributor 档为 $0.10 和 $0.20,条件是允许 Meta 用这些数据训练未来模型。

    推荐理由:文章用两档定价的价差反推客户数据的单价,读者可借此理解基础模型用 token 补贴换训练数据的商业逻辑。

  2. @kimmonismus67

    前沿模型竞争格局在很短时间内从 OpenAI 与 Anthropic 双强之争,扩展为 OpenAI、Anthropic、xAI、Meta 多方并跑、Google 重新加入的局面,中国开源权重模型也紧随其后。

    引用@ArtificialAnlys@ArtificialAnlys

    Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62) Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh) ➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available ➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh) ➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh) Other model details (xhigh variant): ➤ Context window: 1M tokens, unchanged from Muse Spark 1.2 ➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Input modalities: text, image, video ➤ Availability: Meta's first-party API and Muse Code

    推荐理由:作者把前沿模型竞争格局的变化讲清楚,并指出中国开源权重模型已贴近第一梯队,可与2025年的撞墙争论对照。

  3. @rohanpaul_ai67

    Meta 发布 Muse Spark 1.3,称相比 1.2 工具调用减少 20%、生成 token 减少 25%,长编码任务可用更少的模型动作和输出来完成同一交付物。该版本在 MRCR 256K–512K 上得分 98.5,超过 GPT-5.6 Sol;在 JobBench(64.9 对 65.7)、OSWorld(66.9 对 68.3)和 AutomationBench(49.4 对 50.3)上接近 Opus 5。转发的官方公告称该版本当天在 Muse Code 和 API 上线,并预告开放权重版本即将发布。

    引用@finkd@finkd

    Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API. Next up 🍉 and Muse Spark open weights releases coming soon. https://t.co/XQQEDEJGD7

    推荐理由:对比表格列出 Muse Spark 1.3 与 GPT-5.6 Sol、Opus 5 在长上下文和智能体基准上的差距,可看出其性能定位。

9月1日周二
  1. @rohanpaul_ai68

    Meta 的 Muse Code 结束 beta,并开放可编程能力,新 SDK 让开发者可在 Muse 的会话、工具、权限和持久状态之上构建自己的 agent 与应用。工作流可并行协调多个子智能体,跨会话消息让多个相互独立的编码任务彼此协调,而非各自孤立。rewind 功能可回退到会话中较早的安全点,同时撤销该点之后的代码改动和对话状态。

    原始视频预览图;未保存可播放视频URL
    引用@finkd@finkd

    Muse Code is out of beta and now built to handle bigger, more complex engineering tasks. Developers can get started with one command today: curl -fsSL https://t.co/0RApZrEJMv | bash

    推荐理由:原文梳理了 Muse Code 的 SDK、多子智能体并行与 rewind 回退机制,可据此判断其作为编码智能体运行时的协作与容错能力。

8月29日周六
  1. AI前线 · 微信公众号82

    Meta 的 OT 项目被叫停:AI Agent 未能取代员工,事故增四成、工程师救火多七成

    Meta 代号 OT 的组织转型计划曾设想用 AI Agent 接管数千人的日常工作,部分团队削减 60% 人力,但第一轮裁员后内部代码变更同比增长 220% 而真正触达用户的变更仅增 36%,AI Agent 造成的重大技术和安全事故较上一年增加 40%,员工处理这些问题的时间增加多达 70%,扎克伯格在 5 月 19 日晚取消了原定 11 月的第二轮裁员。

    推荐理由:Meta 内部用 Agent 替代员工的转型计划,给出了事故率上升四成、救火时间增加七成等具体量化后果,可供评估 Agent 落地边界。

8月28日周五
  1. InfoQ · 微信公众号79

    Meta 叫停用 Agent 取代员工的 OT 项目,安全事故增四成、工程师救火时间增七成

    Meta 内部代号 OT 的组织转型项目曾设想用 AI Agent 接管数千名员工的日常工作、部分团队裁减 60% 的人,最终因事故频发而叫停。5 月 20 日第一轮裁员落地约 10%,扎克伯格在裁员开始前几小时取消了原定 11 月的第二轮裁员。

    推荐理由:Meta 用 Agent 替代员工的内部转型以事故增加和裁员叫停收场,为评估 Agent 在真实生产环境的可靠性提供了具体案例。

8月25日周二
  1. @Alibaba_Qwen67

    通义千问(Qwen)官方账号转发 natolambert 的分析并致谢,该分析用 Codex 解析了 ChatGPT 发布以来的 50 万篇 arXiv AI/ML 论文。数据显示,2024 年约 30% 论文提及美国开源模型、仅 10% 提及中国模型,如今约 40% 提及中国开源 LLM、25-30% 提及美国模型;提及任一 LLM 的论文中有三分之一提到 Qwen,OpenAI 闭源模型以约 37% 居首,Llama 在 2025 年 4 月达到 30% 峰值后持续下滑。提及 LLM 的论文占比已从 2023 年的 10% 升至 50% 以上。

    引用@natolambert@natolambert

    Over the weekend I had Codex parse 500K arXiv AI/ML papers since ChatGPT to understand which open models are used for research. In 2024, ~30% of papers mentioned an American open model and only 10% a Chinese model. Today, ~40% of papers mention a Chinese (open) LLM, and only 25-30% an American one. Chinese models are the default for research. Chinese mentions are still growing while American open models are stagnating. When looking at this data it's important to remember that papers substantially lag model releases, as research takes a long time. Qwen's steady growth is reflective of this, but so is Llama's lasting power. Some more observations: 1. Qwen has been steadily growing, and today 1/3 of papers which mention any LLM mention qwen. OpenAI's closed models are the highest overall, at ~37%. 2. Llama peaked around April of 2025 at 30% of papers which mention any LLM (including ChatGPT etc). Llama 4 was released at about the same time, and Llama has been declining since. 3. Gemini and Claude are less common than the leading open models, mentioned in 10-15% of papers puts them behind all of Qwen, Llama, and DeepSeek. Open models should be and are the foundations of open research. The % of papers mentioning any LLM have been steadily climbing since 2023. | Year | January | April | July | October | | 2023 | 10.43% | 15.39% | 18.69% | 32.18% | | 2024 | 29.70% | 33.93% | 35.70% | 44.25% | | 2025 | 39.23% | 45.28% | 44.94% | 53.52% | | 2026 | 55.49% | 57.26% | 53.14% | TBD Now over 50% of AI papers, from 10% in 2023. Other notes: - Gemma and Mistral hover around 5-10%. - Our beloved fully-open Olmo models have been ~1% since the first release in Jan. 2024. - DeepSeek has a clear jump after R1 in Jan. 2025 - Data derived from the most popular ML arXiv categories: cs. AI, cs. CL, cs. CV, cs. LG, stat. ML Just like our downloads and derivative model data, this is updated daily on the Interconnects Open Model Dashboard.

    推荐理由:引用数据呈现了近三年论文提及开源模型的份额变化,读者可据此观察中美开源模型在研究社区中的位置。

8月13日周四
  1. 机器之心 · 微信公众号77

    OpenAI、Anthropic、Google 等签署欧盟 AI 内容透明度准则,Claude 公布水印实现方案

    包括 OpenAI、Anthropic、Google、Meta、Microsoft 在内的一批公司签署了欧盟《AI 生成内容透明度行为准则》,承诺推进 AI 生成内容的标记与检测。

    推荐理由:Claude 的水印方案给出了可查验的落地细节,也让人看到文本水印在改写与翻译后可能失效的边界。

8月12日周三
  1. Meta Engineering62

    WhatsApp 推出端侧运行的 Scam Alert 反诈骗功能

    WhatsApp 推出可选功能 Scam Alert,在设备端用机器学习模型对非联系人消息做诈骗分类,消息内容不上传、不自动上报。模型版本经 Cloudflare 签名并发布到第三方 append-only 透明账本,遥测经基于 TEE 的机密联邦分析管线以差分隐私聚合方式回传,用户可在应用内查看 Scam Alert Activity 日志,功能正在 Beta 小范围推送。

    推荐理由:原文给出了 Scam Alert 的端侧架构与可验证机制细节,安全研究者可以据此评估其在端到端加密下的实现。

8月11日周二
  1. 量子位 · 微信公众号78

    Meta 开源 30B 本地 Agent 模型 Muse Glimmer,24GB 显存可运行

    Meta MSL 开源本地 Agent 模型 Muse Glimmer,总参数约 296 亿(含约 18 亿参数视觉编码器),上下文超 13 万 Token,支持文本和图像输入,量化到 17GB 后可装进 24GB 消费级显存。

    推荐理由:Meta 开源 30B 级本地 Agent 模型,量化后 24GB 消费级显存可跑,读者可据此判断本地 Agent 的部署门槛变化。

  2. Manus Blog63

    Manus 致用户信:恢复独立运营,部分用户数据将于 8 月 23 日至 24 日删除

    Manus 宣布恢复独立运营,为遵守监管要求,部分用户在 2025 年 12 月 29 日(Meta 收购日)当天或之后产生的数据将于 2026 年 8 月 23 日 08:00 至 8 月 24 日(SGT)删除。受影响用户可在 8 月 23 日 07:59(SGT)前用备份工具备份数据,8 月 25 日 08:00(SGT)起恢复,期间预计有两天无法访问;未受影响用户无需行动。

    推荐理由:官方说明分拆后数据删除与备份恢复的时间表和操作入口,受影响用户可据此安排备份。

8月10日周一
  1. Hugging Face Blog81

    Meta 发布开源多模态模型 Muse Glimmer-30B,主打本地智能体场景

    Meta 发布从 Muse 蒸馏而来的 30B 多模态模型 Muse Glimmer,采用 Apache 2.0 许可,面向本地隐私场景的智能体用途。模型由 2B ViT 视觉编码器和 28B 文本解码器组成,支持图像、视频、多模态工具调用和目标检测,并附带基于 DFlash 的可选投机解码。

    推荐理由:原文给出架构组成、基准对比和各推理框架的 day-0 用法,读者可以据此评估本地部署的可行路径。

  2. LMSYS Blog60

    SGLang 为 Meta Muse Glimmer 提供 Day-0 支持,面向本地智能体工作流的多模态模型

    SGLang 与 Meta Superintelligence Labs 合作,为 Muse Glimmer 提供 Day-0 支持。Muse Glimmer 是 30B 参数多模态稠密模型,拥有 128k+ token 上下文窗口,架构含 27.9B 文本解码器、1.9B ViT 和多模态 projector。

    推荐理由:原文给出多平台实测吞吐与量化方案,读者可据此评估在本地硬件跑智能体工作流的可行性。

8月2日周日
  1. InfoQ · 微信公众号81

    微软、Meta 同日发布财报:Azure 全年收入破 1000 亿美元,Meta 自由现金流仅剩 7.84 亿美元

    7 月 29 日微软与 Meta 同日发布财报,微软 2026 财年收入 3310 亿美元、Azure 全年收入首次披露突破 1000 亿美元,微软云当季收入 593 亿美元、增长 27%;Meta 第二季度收入 608 亿美元、同比增长 28%,但当季资本开支 310.8 亿美元,自由现金流仅剩 7.84 亿美元、同比下降 91%,股价一度下跌 10%。

    推荐理由:对比同日两份财报的现金流与收入结构,可看到 AI 资本开支回报可见度如何影响市场判断。

7月22日周三
  1. Tomer Tunguz68

    Google Cloud 营收增速与 NVIDIA 趋同

    Google Cloud 2026 年 Q2 营收同比增长 82% 至 248 亿美元,超过 223 亿美元的普遍预期。该分析指出其增速曲线已与 NVIDIA 趋同,云业务积压订单达 5140 亿美元,同比增长 385%,其中略超一半将在 24 个月内转化为收入。Google 本季度开始确认 TPU 系统销售收入,但 CFO 称这部分金额很小,大部分 TPU 收入将在 2027 年入账。

    推荐理由:用财报数字把 Google Cloud 与 NVIDIA 的增长曲线放在一起对比,并给出 AI 基础设施投入的规模参照。

7月15日周三
  1. 量子位 · 微信公众号78

    26名Meta员工起诉公司用AI筛选裁员名单,指其歧视休假员工

    26名Meta现任和前任员工于7月13日在加州北区联邦法院提起集体诉讼,指控Meta用一套内部AI系统筛选裁员名单,对正在休产假、病假或申请残疾合理便利的员工形成系统性歧视。

    推荐理由:诉讼呈现用产出与使用量数据筛选裁员会如何误伤休假员工,为算法决策的审查与可解释性争议提供具体案例。

7月13日周一
  1. IT Home80

    Meta 宣布扩建路易斯安那州数据中心至 5GW,总投资超 500 亿美元

    Meta 官方宣布将位于路易斯安那州里奇兰教区的数据中心算力规模扩大至 5GW,项目投入超过 500 亿美元,用于各项基础设施建设和劳动力培训计划。Meta 承诺该数据中心的能源、水资源及相关基础设施费用全部由其承担,并额外投资超 10 亿美元改善当地道路、供水及污水处理系统。Meta 近期还与安特吉公司达成协议,为七座新建天然气发电厂、三个电网级储能电池、核电增容项目及其他外购电力提供资金支持。

    推荐理由:Meta 将路易斯安那州数据中心扩至 5GW 并投入超 500 亿美元,自建天然气与储能供电的配套值得关注。

6月8日周一
  1. Hugging Face Blog61

    OpenEnv 转为委员会共同治理,成为开源智能体 RL 的互操作层

    OpenEnv 宣布由 Meta-PyTorch、Reflection、Unsloth、Modal、Prime Intellect、Nvidia、Mercor、Fleet AI、Microsoft、Hugging Face 和 RadixArk 组成的委员会共同协调,项目地址迁移至 huggingface/OpenEnv。

    推荐理由:OpenEnv 由多家机构组成的委员会共同协调,并明确为 RL 环境的互操作层,读者可了解开源智能体训练的协作与协议设计。