AI 导读
天哪!Ben 把这个神秘模型跑了 10 个 DeepSWE 任务,得分超过 80%,而 Fable 是 65%,GPT-5.6-sol 是 52%! 这太疯狂了。可能是中国公司。我猜要么是新的 GLM,要么是 Kimi 模型。
正文
Wtf! Ben ran this mystery model through 10 DeepSWE tasks and it scored over 80%, versus 65% for Fable and 52% for GPT-5.6-sol!
this is insane. Probably a chinese company. Either a new GLM or Kimi model, I reckon.
来源:@kimmonismus · x.com