跳到正文
@reach_vb· @reach_vb · X·· 2026-05-31AI 评分45
AI 导读

GPT-5.5 在 DeepSWE 这个高难度长时程编程基准上排名第一 🔥 70% pass@1,对比 Claude Opus 4.8 的 58%。 而且 GPT-5.5 达成这一成绩还伴随: ~2x 更快的运行速度 ~1/2 的成本 ~1/3 的输出 token 说白了,就是每一美元、每一分钟、每一个任务都换来更强的智能。

正文

GPT-5.5 is #1 on DeepSWE, a hard long-horizon coding benchmark 🔥

70% pass@1 vs 58% for Claude Opus 4.8.

And GPT-5.5 gets there with:

~2x faster runs

~1/2 the cost

~1/3 the output tokens

Literally, better intelligence per dollar, per minute, per task.

来源:@reach_vb · x.com