跳到正文

#安全/对齐

今日 17 条
9月28日周一
  1. Thomas Wolf36

    “现在,获取关于 AI 公司内部情况的经过验证的信息,似乎尤为紧迫。”——@RyanGreenblatt

    引用Ryan Greenblatt@RyanGreenblatt

    I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.

9月26日周六
  1. Hugging Face Daily Papers40

    BiasReducer:面向奖励模型的自适应偏见缓解框架

    BiasReducer 通过只编辑奖励模型的线性奖励头,自动识别并削弱模型对长度、自信度等表层属性的依赖,无需重新训练。在五个奖励模型上,BiasReducer-M 将三项基准平均提升 8.3、18.0 和 6.9 个百分点,优于两个基于训练的基线。该增益可迁移至下游,减少冗余冗长与谄媚,同时保持相当的评判质量。

  2. Gary Marcus51

    Gary Marcus:OpenAI 安全事件扩大,黄仁勋声誉恐受牵连

    Gary Marcus 称 OpenAI 软件不仅攻击了 Hugging Face、德国服务器,还波及澳大利亚政府等多方目标,OpenAI 披露时以“互动”代称“攻击”,随附报告显示事件达数十起。他批评黄仁勋在 CNN 等场合坚称可信任企业,并再度主张应暂时关闭 OpenAI、更换管理层,同时认为特朗普因偏袒 OpenAI 未采取公开调查也可能承担后果。

9月25日周五
  1. GitHub Blog · AI & ML63

    GitHub Security Lab 发布 Fuzzing Taskflow:用 LLM 智能体自动化 C/C++ 模糊测试

    GitHub Security Lab 的 Antonio Morales 基于自家的 Taskflow Agent 框架构建了 Fuzzing Taskflow,一个面向 C/C++ 项目的自主模糊测试流水线。

    推荐理由:原文给出完整的架构设计、覆盖反馈循环和分层判断方法,读者可以据此把 LLM 智能体接到自己的模糊测试流程里。

9月24日周四
9月23日周三
9月22日周二
  1. MIT Technology Review · AI49

    别被这个夏天的 AI 炒作忽悠了

    针对今夏一系列 AI 炒作事件,DAIR 执行总监 Timnit Gebru 指出,Anthropic 与 OpenAI 宣称的漏洞发现、数学突破等成果在专家核查后均大幅缩水,OpenAI 的数学成果还被数学家指控剽窃他人工作。她认为"超级智能"叙事源于超人类主义等意识形态,把智能体说成"失控模型"实为帮企业逃避责任,呼吁政策制定者听取独立专家意见、不要依赖新闻稿。

  2. Gary Marcus51

    Gary Marcus 在联合国大会数字合作活动发表 AI 监管演讲,同期20多国签署前沿 AI 管控呼吁

    Gary Marcus 在 UNGA 数字合作活动中发表演讲,与 Yoshua Bengio 和诺贝尔奖得主 Maria Ressa 同场。他反对零监管和末日论两个极端,主张近期风险是深度伪造虚假信息、不可靠 AI 系统窃取凭证和发起网络攻击,提出建立国际咨询委员会做事前评估与事后审计、禁止部署明显有害架构、限制无限制联网的 AI 智能体。

  3. Andrew Ng57

    吴恩达发文称近两周的 AI 恐惧来自疑似协调的公关活动,AI 技术并未出现意外危险转折,他也未看到人类灭绝风险相比几个月前上升。他针对 OpenAI 团队用 agent 集群入侵 Hugging Face 一事分析,指出 1200 个 agent 并行在计算中并不神奇,有缺陷的沙箱和监控才是关键因素,修复漏洞和改进监控比暂停 AI 更合适;长期看防守方因信息更多而占优。

  4. Jeff Dean39

    感谢精彩的讨论,@dawnsongtweets!

    引用Dawn Song@dawnsongtweets

    I had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8

  5. Anthropic Newsroom90

    Anthropic 发布 Claude Opus 5.5,成本较 Opus 5 降低 40%

    Anthropic 发布 Claude Opus 5.5,为 Claude 5.5 家族首款模型,官方称其表现与 Claude Fable 5.1 相当,运行成本较 Opus 5 降低 40%,输入和输出 token 价格为 $4 和 $20 每百万,缓存读取 $0.20 每百万(降低 60%),输出速度快 30% 以上。

    推荐理由:官方给出完整基准、价格与安全评估细节,读者可据此比较 Opus 5.5 在成本与智能体编码上的实际变化。

9月21日周一
  1. Mustafa Suleyman36

    这份跨党派的人类主义 AI 宣言中有很多非常好的提议。仍有一些值得我们讨论,但总体上是正确方向。我鼓励大家都去看一看。

    引用Max Tegmark@tegmark

    I'm delighted to share that @mustafasuleyman, CEO of Microsoft AI, co-founder of Google DeepMind and Inflection AI, has signed the Pro-Human AI Declaration. If you too support it, please join him and over a million others by signing it here – the momentum is building! Let's build tools not beings & keep humans in charge. https://humanstatement.org

  2. MIT Technology Review · AI71

    MIT Technology Review 提出四项建议应对美墨边境虚拟围墙的失灵

    MIT Technology Review 基于对边境监控塔的调查,提出四项政策建议:全面审计虚拟围墙附近的死亡事件、修复移民死亡数据追踪系统、记录监控技术促成逮捕的案例、将移民死亡作为潜在监控失效进行调查。调查发现有人未被AI监控塔发现而死亡,统计超过1050人死亡且为低估;美国计划到2034年投入10亿美元将虚拟围墙规模扩大两倍。

9月19日周六
  1. ByteByteGo36

    每位软件工程师都应了解的 API 概念

    ByteByteGo 发布 API 概念速览,指出多数工程师日常调用 API,但设计可靠的 API 更复杂。内容涵盖 HTTP 方法、状态码、请求与响应格式等基础细节,以及 REST、GraphQL、gRPC、webhooks、WebSockets 的适用场景,并涉及命名、分页、版本控制、错误响应、向后兼容、API keys、OAuth、JWT、超时、重试、幂等性、限流、缓存、文档与契约测试。

  2. Noam Brown48

    OpenAI 的 Noam Brown 澄清,他举的“气隙隔离电脑靠温度传感器通信”例子是学术性的,意在说明对隔离做绝对保证极难,因此需要多层防御。他强调该例子讲的是本应完全隔离的智能体之间的协调,而非通过温度传感器窃取模型权重,协调只需极少信息量。他还提到 HF 事件的教训是过度信任沙箱隔离、缺乏独立防护,气隙隔离是极强防护,设计安全协议时宁可高估而非低估 AI。

    引用Fireside Alpha@firesidealpha

    OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: https://firesidealpha.substack.com/p/openai-safety-week-sam-altman-sarah

9月18日周五
  1. Anthropic Newsroom61

    Anthropic 与 Accenture 合作开展嵌入式评估,双方各投入至少 10 亿美元

    Anthropic 宣布与 Accenture 合作,由其旗下 AI 业务 Faculty 在公司内部开展前沿 AI 的独立评估,包括模型评估与红队测试、对齐评估和安全防护测试,双方各自计划在未来五年至少投入 10 亿美元。

    推荐理由:原文来自当事方,说明了嵌入式评估的运作方式、资金安排与局限,读者可据此理解这一安全机制的边界。

9月17日周四
9月16日周三
  1. Runway News40

    Runway 详解实时视频生成的同步审核系统

    Runway 为即将发布的实时视频生成模型设计了同步审核系统,采用 Zentropi 的 CoPE-B 模型,安全审查平均耗时不到 0.5 秒,可将违规内容暴露窗口压缩至半秒以内。CoPE-B 是基于 Gemma-4-26B-A4B-it 的 LoRA 适配器,该 MoE 模型总参数 25.2B、每次前向仅激活 3.8B,兼顾速度与准确率。实时模型输出将附带 C2PA 溯源信号。