跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-07AI 评分41
AI 导读

研究发现编码模型互相编辑代码时倾向过度修改,而严格提示词无法稳定解决这一问题。研究者用"保持编辑最小但通过构建与测试"两个信号后训练 Olmo3 7B,CROCODIL 将各模型实现的编辑距离大致减半,同时提升构建与全测试通过率。该研究仅限 Rust 函数编辑,但提示混用编码模型的团队应基准测试跨模型编辑并衡量多余 diff 大小。

正文

When coding models edit each other’s work, they tend to over-edit, and this paper shows that training for minimal correct diffs works better than stricter prompts.

A stricter prompt telling the model to make only minimal edits did not fix this consistently.

The researchers instead post-trained Olmo3 7B with 2 signals: keep the edit small, but still build and pass tests. CROCODIL roughly halved its edit distance across implementations from every model, while improving build and all-test pass rates on every foreign implementor tested.

The study is limited to Rust function edits. Still, if a team mixes coding models, it should benchmark cross-model editing and measure unnecessary diff size, not just whether the final code passes.

– arxiv. org/abs/2609.03894

Title: "CROCODIL: Cross-Model Code Editing with LLMs"

来源:@rohanpaul_ai · x.com