跳到正文
原文
@kimmonismus· @kimmonismus · X·· 2026-06-10精选AI 评分68
AI 导读

Anthropic 的 Fable 5 保障机制在前沿 LLM 开发场景下不会直接拒绝或提醒用户,而是通过 prompt modification、steering vectors 和 PEFT 等方式悄悄降低模型自身的有效性。

推荐理由

材料呈现模型在敏感领域被悄然降能的机制,可供观察厂商在能力与安全之间的取舍方式。

正文

Anthropic’s new Fable 5 safeguards are fascinating.

When the model is used for frontier LLM development, it apparently does not simply refuse or warn the user. Instead, it quietly limits its own effectiveness through techniques like prompt modification, steering vectors, and PEFT.

That means Claude may still answer, but become deliberately less useful for building frontier AI systems, pretraining pipelines, distributed training infrastructure, or ML accelerators.

Anthropic says this should affect only around 0.03% of traffic, but the precedent is big: They are being selectively capability-throttled in strategically sensitive domains.

引用NomoreID (@Hangsiin)@Hangsiin
When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model’s capabilities through methods such as prompt modification, steering vectors, and PEFT. Anthropic estimated that this would affect approximately 0.03% of traffic.
在 X 查看被引用的帖子

来源:@kimmonismus · x.com