Dex Horthy 完成了 SlopCodeBench 对 GLM 5.3、Sol 5.6 和 Astra 的完整基准测试,覆盖全部场景(此前仅跑子集),Fable 5.1 结果待公布。Astra 得分略高于 gpt-5.5,但差距不大;Sol 得分低于 gpt-5.5,令他意外。测试历时约一周,因多家服务商中断多次续跑,他建议后续采用多次评估聚合分数的更科学做法。
finally finished a full SlopCodeBench run against GLM 5.3, Sol 5.6, and Astra (Fable 5.1 results coming soon)
This is different from all our previous research on SlopCodeBench (from @GOrlanski @ UW) in that we ran the full benchmark, every single scenario here. Previous runs did a small subset of the challenges.
Asterisks:
- this was run over ~1 week and had to be resumed a few times due to various provider outages
- I still think a more scientific approach here would be to do what's common with other benchmarks, which is run multiple evaluations and aggregate the scores
Looks like Astra scores a few points higher than gpt-5.5 here. Not as many as I'd think. I'm surprised Sol got lower than gpt 5.5 because in my experience I like working with Sol a bit more.
My vibes-best guess is that the newer models are likely to go more off the rails on higher thinking modes (e.g. I almost always use Sol in medium or low effort for most work on @humanlayer_dev)
来源:@dexhorthy · x.com