跳到正文
DAIR.AI· @dair_ai · X·· 2 小时前AI 评分52
AI 导读

Meta Superintelligence Labs 等发布 MIMESIS,一个基于真实人类对话和 13 种行为模式训练的 9B 用户模拟器,解决用助手 LLM 模拟用户时过于配合和explicit的问题。

正文

Banger paper from Meta Superintelligence Labs on user simulators for agent training.

Agent RL setups usually let an assistant LLM play the user, so the simulated user is too cooperative and too explicit.

A fixed GPT-5.5 agent finds tau-bench tasks easier with these users than with real people.

This work trains MIMESIS, a 9B user simulator, on human conversations and 13 behavior patterns observed in real users. It beats Claude Opus 5 on behavioral fidelity by 13.4 points.

Agents trained against it outperform agents trained against GPT-5.5 under all nine user simulators they never saw.

They find that adding a coaching step that turns the simulator's private reasoning into feedback adds further gains.

Paper: https://academy.dair.ai/papers/mimesis-learning-user-simulators-as-training-environments-for-interactive-agents-2610.09484

来源:DAIR.AI · x.com