Elvis Saravia 指出,用测试框架(harness)衡量模型质量的方式已完全失效,他更倾向用 Pi、Hermes Agent 等极简框架测试模型。他认为行业缺少标准化评测方式,而框架工程正是领先 AI 公司投入的重点。Claude 模型已能在一定程度上按任务动态生成框架,但表现不稳定且机制不明,这意味着框架可能像系统提示词一样只是可调产物,benchmark 评测将愈发模糊。
Important discussion. Measuring models against harnesses is completely broken. I prefer to test model quality against minimal harnesses like Pi and Hermes Agent. This is not perfect, as there are biases in the harnesses that favor some models and not others.
A standard way to do this is missing but important, as harness engineering is where leading AI companies are focusing efforts. Not enough effort here, as things are moving fast and harness optimization is individualistic.
On the flip side, I feel like models will eventually have the ability to dynamically generate harnesses on the fly as per task. Claude models do this already to some extent, though pretty inconsistently and remain a mystery. But this could mean that a harness is just a tunable artifact like a system prompt. In that realm, how are we assessing it, and exactly what? Benchmarking will only get murkier from here onwards.
来源:@omarsar0 · x.com