研究发现,模型编辑其他模型写的代码时改动明显多于编辑自己写的代码,不同训练数据带来的风格偏好是原因之一。CROCODIL 后训练框架用相似度奖励惩罚大改动、用执行奖励评估构建与测试是否通过,两者相乘而非相加,从而避免策略靠让任务失败来减少改动。
This is a weird behavior in coding models and something worth looking into.
It turns that some models over-edit code that another models wrote.
There is a high chance that your repo now has commits from more than one model, and that changes how each of them edits.
Researchers measured what happens when one model edits code another model wrote. Different training data produces different stylistic preferences, and models make more edits, often excessive ones, on foreign code than on their own.
CROCODIL is a post-training framework that reduces that behavior. A similarity reward penalizes large changes and an execution reward scores build and test success, and the two are multiplied rather than added. That product stops the policy from shrinking edits by simply failing the task.
Paper: https://t.co/WyzDs0YWiO
来源:@omarsar0 · x.com