AI 导读
(4/6) 跨框架泛化:在 QwenClawBench 和 CoWorkBench 上,无论评估时使用何种框架,Qwen3.7-Max 都表现出强劲且一致的性能,证实该模型学到的是解决任务的能力——而非利用特定框架的漏洞。
正文
(4/6) Cross-Harness Generalization: Across QwenClawBench and CoWorkBench, Qwen3.7-Max delivers strong, consistent performance regardless of the harness used at evaluation time, confirming that the model has learned to solve tasks — not to exploit particular harnesses.
来源:@alibaba_cloud · x.com