AI 导读
在我们的各项基准测试中,该模型树立了新标准。 它在 Terminal-Bench-Science 0.1 上得分 52.6%,是 Fable 5 的两倍多。在 Terminal-Bench 4.0 上,它得分 55.8%,而 Fable 5 为 42.0%。https://t.co/aSb72LSxee
正文
Across our benchmarks, the model sets a new standard.
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. https://t.co/aSb72LSxee
来源:@claudeai · x.com