跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-07AI 评分51
AI 导读

微软一篇新论文提出把测试时推理成本摊销为蒸馏技能:收集 35–50 条历史轨迹,由编码智能体提取反复出现的失败模式,再将其写成 markdown 技能加入非推理模型的系统提示词。

正文

What if you could pay the reasoning cost once, then reuse what the model learned across future tasks?

New Microsoft paper finds that some expensive test-time reasoning can be replaced with a small set of rules learned from previous agent runs.

The paper tests a cheaper alternative: collect 35–50 past trajectories, have a coding agent extract recurring failure patterns, then turn those patterns into a small markdown skill added to the non-reasoning model’s system prompt.

For GPT-5.4-mini, those skills recovered 55%–100%+ of the gap between non-reasoning and reasoning modes across 4 agent benchmarks, while using 2.9–4.5× fewer output tokens than reasoning.

On ALFWorld and τ²-retail, the skilled non-reasoning model actually beat the reasoning mode.

The useful part is that the distiller did not need reasoning traces: skills built only from cheap non-reasoning rollouts were competitive across all 4 domains.

The limit is equally useful: reasoning still won on telecom and SpreadsheetBench, where each task contains more instance-specific dependencies that a fixed skill cannot capture.

So the practical split is: distill repeated procedures once, then reserve expensive test-time reasoning for the tasks that genuinely need fresh search.

来源:@rohanpaul_ai · x.com