跳到正文
@kimmonismus· @kimmonismus · X·· 22 天前AI 评分43
AI 导读

又来了:一个神秘新模型在 DeepSWE 上得分 74%,击败 GPT-5.6 Sol 和大多数其他模型。 据称它在 Terminal-Bench 2.1 和 SWE-Bench Verified 上也优于 GPT-5.6 Sol,同时价格便宜得多。我赌是又一个中国模型。

正文

Here we go again: a new stealth model scores 74% on DeepSWE, beating GPT-5.6 Sol and most other models.

It also reportedly outperforms GPT-5.6 Sol on Terminal-Bench 2.1 and SWE-Bench Verified, and being much cheaper at the same time. My bet is on another chinese model. https://t.co/TT0OyTRZoE https://t.co/sGAgfYyRqe

来源:@kimmonismus · x.com