跳到正文
原文
ByteByteGo· ByteByteGo·· 2026-08-25精选AI 评分68

研究者用同族小模型还原 Anthropic、OpenAI 与 Google 的加密推理块

How to Steal an AI Model’s Private Thoughts

AI 导读

2026 年 8 月,MATS Research、ELLIS Institute Tübingen 与 Max Planck Institute for Intelligent Systems 的研究者测试了 Anthropic、OpenAI 和 Google 返回给客户端的加密推理块,发现把块重放进同族更便宜的模型,就能让隐藏推理以明文输出。

推荐理由

研究揭示加密推理块可被同族小模型复述,公开 agent 日志里的密钥无法靠清洗明文保护。

来源:ByteByteGo · blog.bytebytego.com