AI 导读
Opus 4.8 在 DeepSWE 上相比 Opus 4.7 有明显提升,同时降低了每任务的平均成本。 不过,GPT-5.5 xhigh 仍以相当明显的优势胜过它,而且更便宜。 OpenAI 最近在模型上真是拼得厉害。非常期待 GPT-5.6 会带来什么。 话虽如此,我得承认:我也开始真的很喜欢 Opus 4.8 了。 我们已经进入了一个两家前沿实验室都在持续发布真正令人印象深刻的模型的时代。
正文
Opus 4.8 is a solid jump over Opus 4.7 on DeepSWE, while also lowering the average cost per task.
However, GPT-5.5 xhigh still beats it by a pretty clear margin while being cheaper.
OpenAI has been cooking insanely hard with its models lately. Really excited to see what GPT-5.6 brings.
That said, I have to admit: I’m starting to really like Opus 4.8 as well.
We’ve entered a moment where both frontier labs keep shipping genuinely impressive models.
Opus 4.8 is now on DeepSWE. On the default high thinking effort, it scores 6% higher than Opus 4.7 xhigh, while also lowering average cost per task. Video在 X 查看被引用的帖子
来源:@kimmonismus · x.com