研究发现编码模型互相编辑代码时倾向过度修改,而严格提示词无法稳定解决这一问题。研究者用"保持编辑最小但通过构建与测试"两个信号后训练 Olmo3 7B,CROCODIL 将各模型实现的编辑距离大致减半,同时提升构建与全测试通过率。该研究仅限 Rust 函数编辑,但提示混用编码模型的团队应基准测试跨模型编辑并衡量多余 diff 大小。
When coding models edit each other’s work, they tend to over-edit, and this paper shows that training for minimal correct diffs works better than stricter prompts.
A stricter prompt telling the model to make only minimal edits did not fix this consistently.
The researchers instead post-trained Olmo3 7B with 2 signals: keep the edit small, but still build and pass tests. CROCODIL roughly halved its edit distance across implementations from every model, while improving build and all-test pass rates on every foreign implementor tested.
The study is limited to Rust function edits. Still, if a team mixes coding models, it should benchmark cross-model editing and measure unnecessary diff size, not just whether the final code passes.
– arxiv. org/abs/2609.03894
Title: "CROCODIL: Cross-Model Code Editing with LLMs"
来源:@rohanpaul_ai · x.com