跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-01AI 评分37
AI 导读

论文提出 CREST,为多轮智能体的每一轮单独打分并给予可验证的信用,再用同一模型作为自教师,把更多学习权重放在该轮中不确定的决策上,但教师只能增强更新、不能推翻验证器的判断。在 Qwen3-4B-Instruct 上,BFCL V3 平均准确率达 52.0%,高于最强 RL 基线的 49.25%。

正文

This paper shows a better way to train multi-turn agents:

score each turn separately, then use a self-teacher to focus learning without letting it override the reward.

Standard RL has a basic problem.

A long agent session can contain successful and failed turns, yet 1 overall reward can blur them together.

CREST fixes that by giving each turn its own verified credit, then using the same model as a teacher to put more learning weight on uncertain decisions inside that turn.

The teacher can strengthen an update, but it cannot reverse the verifier's judgment.

On Qwen3-4B-Instruct, it reaches 52.0% average BFCL V3 accuracy versus 49.25% for the strongest RL baseline.

– arxiv. org/abs/2608.13179

Title: "Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents"

来源:@rohanpaul_ai · x.com