AI 导读
东京都立大学的新论文发现,AI 文本检测器 Pangram 漏检 79.8% 经 Meta Muse-Glimmer 改写的科学摘要,但对 GPT-5 改写摘要的检出率为 93.5%,且在 5,000 篇人写摘要中仅误报 1 篇(假阳性率 0.02%)。漏检率主要取决于执行改写的 LLM 版本,两者 Spearman 相关性 ρ=0.25 不显著;按学科看,漏检分布在健康科学为 45.2%、生命科学 33.4%、物理科学 31.4%、社会科学 26.5%。
正文
Pangram, the AI-text detector, missed 79.8% of scientific abstracts rewritten by Meta's Muse-Glimmer, while flagging just 1 of 5,000 human abstracts.
In a new paper from Tokyo Metropolitan University, reseaerchers find the share of AI-rewritten abstracts that Pangram misses depends strongly on the LLM version
Shows that it caught 93.5% of GPT-5 rewrites but missed 79.8% from another new model. Its miss rate depended mostly on which model did the rewriting.
来源:Rohan Paul · x.com