跳到正文
@op7418· @op7418 · X·· 2026-06-11AI 评分62
AI 导读

Anthropic 宣布调整 Fable 5 的安全护栏,本周起被标记的请求会明显降级到 Opus 4.8,API 端还将返回拒绝理由。Anthropic 表示此前为快速上线而采用不可见护栏是错误的权衡,用户理应看到护栏的存在与原因。由于可见护栏更易被绕过,Anthropic 预计需加强对越狱的鲁棒性,期间会出现更多误判,同时也在调整生物与网络安全分类器以减少对无害请求的触发。

正文

Anthropic 滑轨,承诺未来几天会在 fable 5 降级到 Opus 4.8 的时候提供明显的通知。

引用ClaudeDevs (@ClaudeDevs)@ClaudeDevs
We’re rolling out changes to make Fable 5’s safeguards for frontier LLM development visible. Starting this week, flagged requests will visibly fall back to Opus 4.8—the same as our safeguards for cyber and bio. You will see this every time it happens. On the API, any flagged requests will return a reason for their refusal (coming to server-side fallback in the next few days). We wanted to deploy Fable 5 to our users quickly and safely. Visible safeguards can be probed, so they have to be robust, which takes time to get right. Invisible safeguards can be targeted more narrowly, allowing us to ship quickly with very few false positives. We went with invisible safeguards for this reason—and that was the wrong tradeoff. You should have visibility into the safeguards we have in place, and why. We’re sorry for not getting the balance right. Making the safeguards visible makes them easier to work around, so keeping them robust to jailbreaks will unfortunately mean more false positives while we improve the classifiers. We're also tuning our bio and cyber classifiers to trigger less often on harmless requests. We know this is frustrating and we’ll do our best to keep this period as short as possible. If you think a request has been mistakenly flagged: run /feedback in Claude Code, click thumbs-down on the fallback in Claude.ai or Cowork, or file the safeguard appeal form for API requests. Your reports help us tune these classifiers and we appreciate your feedback. support.claude.com/en/articl…
在 X 查看被引用的帖子

来源:@op7418 · x.com