发表机构
Meta Superintelligence Labs; New York University; University of Wisconsin - Madison(Meta超级智能实验室; 纽约大学; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MIMESIS通过训练用户模拟器作为交互环境,提升智能体训练效果,其9B模型在行为保真度上超越前沿模型,并借助CSD方法进一步强化智能体性能。
AI 中文摘要
训练和评估交互式语言智能体通常需要丰富的用户交互,然而收集人类反馈成本高昂且难以扩展。模拟用户提供了一种可扩展的替代方案,但它们必须既类似于真实用户行为,又为智能体提供有用的学习经验。相比之下,大多数智能体训练框架依赖于现成的助手型大语言模型,其乐于助人的特性可能使其与真实用户相比过于合作、过于明确且行为同质化。我们提出了MIMESIS,一个专门构建的用户模拟器,它在人类对话上进行训练,并带有显式的推理监督和从真实用户交互中提炼出的13种现实行为模式。实验上,我们的9B模型达到了65.7的SOUL指数,超过了最强的前沿模型。与Claude-Opus-5(在RealUserSim和SimulatorArena上的最强基线)相比,MIMESIS在行为保真度上提高了13.4个百分点,在图灵距离上分别降低了3.6个百分点。然后,我们冻结模拟器,并通过与冻结模拟器进行多轮强化学习来训练智能体。在八个环境中,使用MIMESIS训练得到的智能体性能优于使用GPT-5.5训练得到的智能体,在所有九个未见过的用户模拟器上均表现更好,展示了对新用户模拟器更强的泛化能力。此外,我们提出了监督式在线自蒸馏(CSD),它利用模拟器生成的私有推理轨迹和后续话语作为关于智能体在多大程度上满足用户需求的反馈。一个教练将这些信息转化为简洁的指导笔记,描述智能体如何在交互过程中更好地预测用户需求并调整其行为。CSD将这种反馈转化为密集的、令牌级别的监督,超越了稀疏的任务奖励,在全部九个评估用户模型上带来了进一步的性能提升。
英文摘要
Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.
CommentsProject page: https://viethoang1512.github.io/mimesis/