跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-20AI 评分46
AI 导读

论文发现恶意检索文档会提升 token 置信度与输出一致性,使基于不确定性的检测器失效。攻击下注意力会集中到投毒文档而非分散于检索证据,作者称之为"注意力坍缩"。因此仅检查最终答案或置信度可能漏报,监控注意力在检索文档间的分布或可提前暴露投毒。

正文

RAG poisoning can make a model more confident, which is exactly why confidence-based detectors can fail.

This paper finds malicious retrieved documents can increase token confidence and output consistency.

So uncertainty-based detectors can miss the attack because poisoning can create false confidence.

Under attack, attention becomes concentrated on poisoned documents instead of staying spread across the retrieved evidence.

The authors call this Attention Collapse.

So for RAG security, checking only the final answer or its confidence may miss the warning.

Monitoring how attention is distributed across retrieved documents could expose poisoning before the answer visibly breaks.

– arxiv. org/abs/2608.06947

Title: "When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse"

来源:@rohanpaul_ai · x.com