跳到正文
@kimmonismus· @kimmonismus · X·· 2026-08-21AI 评分22
AI 导读

天哪!Ben 把这个神秘模型跑了 10 个 DeepSWE 任务,得分超过 80%,而 Fable 是 65%,GPT-5.6-sol 是 52%! 这太疯狂了。可能是中国公司。我猜要么是新的 GLM,要么是 Kimi 模型。

正文

Wtf! Ben ran this mystery model through 10 DeepSWE tasks and it scored over 80%, versus 65% for Fable and 52% for GPT-5.6-sol!

this is insane. Probably a chinese company. Either a new GLM or Kimi model, I reckon.

来源:@kimmonismus · x.com