跳到正文
@alibaba_cloud· @alibaba_cloud · X·· 2026-05-21AI 评分44
AI 导读

(4/6) 跨框架泛化:在 QwenClawBench 和 CoWorkBench 上,无论评估时使用何种框架,Qwen3.7-Max 都表现出强劲且一致的性能,证实该模型学到的是解决任务的能力——而非利用特定框架的漏洞。

正文

(4/6) Cross-Harness Generalization: Across QwenClawBench and CoWorkBench, Qwen3.7-Max delivers strong, consistent performance regardless of the harness used at evaluation time, confirming that the model has learned to solve tasks — not to exploit particular harnesses.

来源:@alibaba_cloud · x.com