AI 导读
让一个问题出错有两种方式:漏掉一个真实的错误,或者把一个没问题的构建标记为错误。许多近期前沿模型偏向其中一种失败模式。例如,Gemini 3.1 Pro 几乎不通过任何正确的构建,而 GPT-5.4 几乎抓不到任何错误。https://t.co/F8K7zkQ5DF
正文
There are two ways to get a question wrong: miss a real mistake, or flag a build that's fine. Many recent frontier models favor one failure mode over another. Gemini 3.1 Pro, for example, passes almost no correct builds, whereas GPT-5.4 catches almost no mistakes. https://t.co/F8K7zkQ5DF
来源:@EpochAIResearch · x.com