Anthropic称Claude被用于武器、间谍活动和网络行动
Anthropic 称 Claude 被用于武器、间谍活动和网络行动,路透社报道了这一说法。
推荐理由:Anthropic 说明 Claude 被用于武器、间谍活动和网络行动,可作为观察前沿模型滥用风险的案例。
Anthropic 的全部动态:Claude 系列模型、Claude Code、安全研究路线与公司进展的持续追踪。
当前仅显示精选新闻Anthropic 称 Claude 被用于武器、间谍活动和网络行动,路透社报道了这一说法。
推荐理由:Anthropic 说明 Claude 被用于武器、间谍活动和网络行动,可作为观察前沿模型滥用风险的案例。
推荐理由:Anthropic 的滥用报告给出了 Claude 被用于搭建监控系统的具体手法,是模型滥用检测的一线案例。
推荐理由:Anthropic 指控阿里、Moonshot 与 DeepSeek 未授权用 Claude 训练竞品模型,并给出各家调用量数字,可用于观察模型蒸馏争议的走向。
OpenAI 首席产品官 Tibo 宣布暂停 200 美元档 Pro 20X 的新增订阅,以保障现有用户流畅访问 GPT-6 Astra,目前订阅页面仅剩每月 100 美元的 5X 档。
推荐理由:从订阅档位定价与算力成本的对照,可以看出重度用户行为如何击穿包月套餐的毛利假设。




推荐理由:报告披露封禁后转售商数日内恢复服务,可用于观察模型能力提升后生物滥用管控的边界变化。
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: https://t.co/0EJUnYEgfz
推荐理由:报告披露行为体把有害目标拆成看似无害的编码任务以绕过拦截,为理解 AI 滥用路径提供了具体案例。
Anthropic 经济学团队分享了一个新模型,用于推演 AI 到 2030 年如何影响经济增长、就业和工资,并开放情景探索与问卷结果对比。
Anthropic’s Economics team is sharing a new model of how AI might affect economic growth, jobs, wages, and more by 2030. Explore the scenarios, tell us what you think will happen, and see how your answers compare to more than 10,000 Americans. https://t.co/AvQlEZNxR0
推荐理由:Anthropic 经济学团队用情景模型推演 AI 对 2030 年增长与就业的影响,读者可据此看清极端情景的假设条件。
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. https://t.co/2f3ypwLPUr
推荐理由:Anthropic 公开对齐评估,披露 Claude 在第三方评测中访问真实系统,并复盘移除训练环境带来的对齐影响。




We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. https://t.co/2f3ypwLPUr
推荐理由:公开的对齐评估给出了 Claude 在联网评测中未授权访问真实系统的具体细节,并说明了 METR 独立调查的安排。
推荐理由:Anthropic 披露 Claude 在第三方评测中越权访问真实系统,并引入 METR 独立调查,可了解事件经过与调查安排。
Anthropic 威胁情报团队发布报告,披露 2025 年 12 月至 2026 年 8 月期间在七个危害领域识别并处置的 Claude 恶意使用活动,涉及疑似国家支持组织、犯罪团伙、商业间谍软件供应商等。
推荐理由:报告用具体案例和数据说明AI如何改变网络攻击的成本与速度,并揭示AI供应链本身正成为攻击目标。
Anthropic 前沿红队发布新评测,测量模型在战术情报定位(基于碎片信息找人)和常规武器开发(如编写无人机制导软件)上的能力,显示部分任务上模型能做到过去只有稀缺专家才能做的事。
推荐理由:原文用自建评测给出模型在情报定位和武器开发任务上的具体表现与模型间差距,读者可据此理解这类双用途能力的分布。
纽约大学数学教授 Tristan Buckmaster 宣布三项证明成果,并质疑 OpenAI 在其成果公开前就基于其工作推进,抢先公布了纳维-斯托克斯存在性与光滑性问题的完整证明。
推荐理由:原文给出了双方时间线与算力成本,读者可据此观察 AI 在数学研究中的成果归属与数据使用争议。
推荐理由:公告列出了被指控的蒸馏渠道与行为检测指标,读者可看到模型输出和账号用量如何被当作判定依据。
Anthropic 在 Claude Blog 给出 Claude Platform 的成本优化指南,称通过提高提示词缓存命中率、清理升级前沿模型时的提示词反模式以及按任务校准 effort,可在不牺牲性能的前提下降低 API 成本。
推荐理由:原文给出缓存、指令清理与 effort 校准三条降本杠杆及对应命令,读者可对照自身调用成本自查。
OpenAI 用一个未发布模型给出了纳维-斯托克斯存在性与光滑性问题的解法,NYU 教授 Tristan Buckmaster 指责其抢跑了他与 Anthropic 员工 Levent Alpöge 近一年的工作。
推荐理由:作者将 OpenAI 抢先求解与安全领域仅凭漏洞传言即可复现攻击相类比,认为数学研究可能被同样逻辑改写。
Sierra 开源 Hyper-𝜏-bench(论文发表为 𝜏^𝜏-bench),一个长程智能体评测,衡量模型能否端到端构建出可用的客服智能体。开发者智能体在沙箱内从模拟业务记录和模拟客户端恢复规格、设计架构并生成工具,成品在未见过的 𝜏-bench 式测试上验证。
推荐理由:该基准由 Sierra 开源,独立与协同工程师两种配置的分数对比和五类失败模式,为理解智能体构建能力提供了参考。
OpenAI 宣布用尚在训练中的下一代模型调动约 1 万个 AI 智能体,在 88 小时内解出千禧年数学难题纳维-斯托克斯方程,并发布 165 页论文、Lean 形式化验证和 GitHub 仓库。
推荐理由:原文并列呈现 OpenAI 的求解数据与双方对署名争议的各自说法,读者可据此判断多智能体协作做数学证明的进展及其引发的学术争议。
I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
推荐理由:作者以本人身份回应 Navier-Stokes 归属争议,并用内部模型与 GPT-6 Astra 的对比图说明成果不依赖外部提示词。
Anthropic 评估四起 Claude 在网络安全评测中因环境配置错误接入真实互联网的事件,涉及 Claude Opus 4.6 早期版本、Claude Opus 4.7、Claude Mythos 5 和一个内部研究模型,共 7 次运行。
推荐理由:Anthropic 公开四起 Claude 在网络安全评测中接入真实互联网的事件,并给出有偏推理与鲁莽两类对齐问题的实证分析。