跳到正文
elvis· @omarsar0 · X·· 3 天前AI 评分59
AI 导读

Meta Superintelligence Labs 论文发现,预训练 LLM 配上轻量推理框架后可在足够测试时预算下超越 RL 后训练版本。后训练把任务推向必解或必不解,提升一致性但降低覆盖,作者称之为 Sharpening Tax;在 14 类基座/后训练配对、42 组设置中普遍存在,其 PTGS 方法按提示词难度自适应调整 RL 采样温度,付出更小代价并提升 pass@1。论文见 https://academy.dair.ai/papers/sharpening-tax-in-post-training-2610.01509。

正文

Great paper for all you pre-training and post-training nerds. The surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents and can even surpass post-trained counterparts with a sufficient test-time budget.

引用DAIR.AI@dair_ai
Banger paper from Meta Superintelligence Labs. They find something super interesting and unexpected. (bookmark it) Base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples. Post-trained models win on pass@1. At large K, base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop. This is because post-training pushes each task toward always solved or never solved. Consistency goes up, and coverage goes down. The authors call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it shows up in most settings, grows with model size, and can be estimated from a few rollouts. Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass@1. Paper: https://academy.dair.ai/papers/sharpening-tax-in-post-training-2610.01509
在 X 查看被引用的帖子

来源:elvis · x.com