AI 导读
小米直播了 V2.6 背后的 RL 训练过程:每步约 2B tokens,1568 条提示词 x 16 次 rollout,完全异步,跨混合 harness 的多任务智能体 RL。细节将在未来几周内开源。 @_LuoFuli 讲述他们如何扩展:https://t.co/2R8YXiVO4m
正文
Xiaomi live-streamed the RL run behind V2.6 while it trained: ~2B tokens per step, 1568 prompts x 16 rollouts, fully async, multi-task agentic RL across mixed harnesses. Details are being open-sourced over the coming weeks.
@_LuoFuli on how they scaled it: https://t.co/2R8YXiVO4m
来源:@OpenRouter · x.com