跳到正文
@EpochAIResearch· @EpochAIResearch · X·· 14 天前AI 评分34
AI 导读

让一个问题出错有两种方式:漏掉一个真实的错误,或者把一个没问题的构建标记为错误。许多近期前沿模型偏向其中一种失败模式。例如,Gemini 3.1 Pro 几乎不通过任何正确的构建,而 GPT-5.4 几乎抓不到任何错误。https://t.co/F8K7zkQ5DF

正文

There are two ways to get a question wrong: miss a real mistake, or flag a build that's fine. Many recent frontier models favor one failure mode over another. Gemini 3.1 Pro, for example, passes almost no correct builds, whereas GPT-5.4 catches almost no mistakes. https://t.co/F8K7zkQ5DF

来源:@EpochAIResearch · x.com