跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 2026-09-02AI 评分30
AI 导读

Claude Fable 5.1(max)尝试回答了 93.4% 的 AA-Omniscience 问题,而 Claude Opus 5 为 87.8%,并录得我们测得的最高准确率 67.2%,领先于 Claude Fable 5 的 65.4%。 在更高的尝试率下,它的幻觉也更多:在它未答对的问题中,它尝试作答的比例为 72.6%,而 Claude Fable 5 为 63.6%。两种效应相互抵消,其 AA-Omniscience Index 得分与 Claude Fable 5 持平。

正文

Claude Fable 5.1 (max) attempts 93.4% of AA-Omniscience questions against Claude Opus 5's 87.8%, and records the highest accuracy we have measured at 67.2%, ahead of Claude Fable 5 at 65.4%.

It hallucinates more with this higher attempt rate: of questions it didn’t get correct, it attempted to respond in 72.6% against Claude Fable 5's 63.6%. The two effects cancel out, and its AA-Omniscience Index score is level with Claude Fable 5.

来源:@ArtificialAnlys · x.com