跳到正文

#安全/对齐

今日 18 条
9月16日周三
9月15日周二
9月14日周一
  1. Mustafa Suleyman40

    这是一个非常直白且符合常识的观点:技术的目的是服务人类,加速人类繁荣。 任何无法实现这一目标的技术都是失败的,应当被拒绝。 我们还没有到那一步。但开始为这种可能性做准备是正确的。

    引用Satya Nadella@satyanadella

    Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive. And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control. So, in this context, we welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like "embedded evaluators" and the broader efforts to develop the mechanisms to make this more than just talk. The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia. This is the approach we are taking: broad access and choice at every layer of the AI stack; enterprise control of learning loops and models; and the “Code of Conduct” that underlies our own first party MAI models that we’ll publish tomorrow for public consultation.

9月13日周日
  1. Demis Hassabis62

    Demis Hassabis 发文表示 Dario Amodei 新文《We Must Pace the Frontier》指出了正确的前进方向,细节仍需完善,但方向对应对这一关键时刻是正确的。他还附上自己此前提出的前沿 AI 行业标准机构提案链接;Dario 原文宣布 Anthropic 将为第三方评估者提供永久员工级系统访问权限。

    引用Dario Amodei@DarioAmodei

    We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier

  2. Peter McCrory69

    Dario Amodei 发表文章《We Must Pace the Frontier》,主张 AI 行业应放慢速度并给出三部分计划,Anthropic 单方面承诺执行其中第一步。该步骤是向第三方评估者提供永久的员工级系统访问权限,用于核验安全措施落实、报告事故并评估训练中模型的对齐情况。全文见 https://darioamodei.com/post/we-must-pace-the-frontier,作者 McCrory 认为嵌入式评估者是合理的第一步。

    引用Dario Amodei@DarioAmodei

    We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier

    推荐理由:Anthropic 首席经济学家推荐 Dario Amodei 新文,提出给第三方评估者永久员工级访问权以核验安全措施,可了解行业自律的具体动作。

  3. Aidan Gomez47

    卡特尔这边有些“好主意”: - 你们得给我们员工级别的权限,访问你们整个运营体系 - 如果我们觉得你们不够“安全”,抱歉,为了“安全”我们得把你们关停 - 中国不会遵守,但其他所有人都得遵守!不然就没芯片! 真是绝了。

    引用Sam Altman@sama

    I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.

9月10日周四
  1. OpenAI News61

    Paul Christiano 加入 OpenAI 基金会董事会

    OpenAI 宣布 Paul Christiano 加入 OpenAI 基金会董事会,并进入其安全与安全委员会。公告称他将带来在 AI 对齐、安全和标准方面的经验。

    推荐理由:官方宣布 Paul Christiano 加入 OpenAI 基金会董事会及其安全与安全委员会,人事动向本身即读者可关注的事实。

  2. Anthropic Newsroom84

    Anthropic 发布 2026 年 9 月威胁情报报告,披露多起 Claude 恶意使用案例

    Anthropic 威胁情报团队发布报告,披露 2025 年 12 月至 2026 年 8 月期间在七个危害领域识别并处置的 Claude 恶意使用活动,涉及疑似国家支持组织、犯罪团伙、商业间谍软件供应商等。

    推荐理由:报告用具体案例和数据说明AI如何改变网络攻击的成本与速度,并揭示AI供应链本身正成为攻击目标。

  3. Anthropic Research74

    Anthropic 红队发布 AI 模型战术情报定位与常规武器能力评测报告

    Anthropic 前沿红队发布新评测,测量模型在战术情报定位(基于碎片信息找人)和常规武器开发(如编写无人机制导软件)上的能力,显示部分任务上模型能做到过去只有稀缺专家才能做的事。

    推荐理由:原文用自建评测给出模型在情报定位和武器开发任务上的具体表现与模型间差距,读者可据此理解这类双用途能力的分布。

9月9日周三
  1. Anthropic Research81

    Anthropic 发布四起 Claude 网络安全评测中误连真实互联网事件的对齐评估

    Anthropic 评估四起 Claude 在网络安全评测中因环境配置错误接入真实互联网的事件,涉及 Claude Opus 4.6 早期版本、Claude Opus 4.7、Claude Mythos 5 和一个内部研究模型,共 7 次运行。

    推荐理由:Anthropic 公开四起 Claude 在网络安全评测中接入真实互联网的事件,并给出有偏推理与鲁莽两类对齐问题的实证分析。

9月7日周一
9月4日周五