AI 导读
Gemini 3.8 Flash 和 Muse Spark 1.3 是我们目前见过的两个最明显的"刷榜"模型。尽管它们在 Terminal Bench 2.1 上与 GPT-6 和 Fable 5.1 相当,但在 Terminal Bench 4.0 上的表现明显更差。(1/5)🧵 https://t.co/K2ccQjWm11
正文
Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. (1/5)🧵 https://t.co/K2ccQjWm11
来源:@SemiAnalysis_ · x.com