发表机构
Peking University; Kuaishou Technology; Institute of Software, Chinese Academy of Sciences; Institute of Information Engineering; Beijing University of Posts and Telecommunications(北京大学; 快手科技; 中国科学院软件研究所; 信息工程研究所; 北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TRACER多轮用户模拟器,通过两阶段训练(监督微调与多轮强化学习)显式建模用户意图演变,在真实客服会话上超越基线11.4转换F1,并引入动态营销基准揭示响应质量与转换率非正相关。
AI 中文摘要
忠实的用户模拟对于大规模构建、评估和改进交互式AI至关重要。然而,看似合理的个体响应并不能确保模拟用户能够复现真实交互中观察到的意图演变和结果。我们提出了TRACER,一种多轮用户模拟器,它显式地建模用户不断演变的意图,并学习将模拟行为与真实交互轨迹对齐。TRACER通过两个阶段进行训练:首先在真实用户对话上进行监督微调,然后进行多轮强化学习。强化学习阶段将分层的结果级和轨迹级奖励与偏差感知的优势调节相结合,共同缓解长对话中的奖励稀疏性和信用分配问题。在按参考队列组织的真实客户服务会话中,TRACER-7B在转换F1分数上超越了最强基线11.4个百分点,同时实现了最低的组级转换率误差和语义轨迹距离,并能够泛化到分布外场景。人类图灵测试产生的识别准确率接近随机水平,支持了生成对话的感知自然性。基于此模拟器,我们进一步引入了动态营销基准,通过模拟交互联合评估LLM的说服效果和响应质量,揭示出更高的响应质量并不一定对应更高的转换率。
英文摘要
Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. Yet current simulators often produce plausible individual responses without reproducing the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that models evolving user intent and aligns simulated trajectories with real ones. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly addressing reward sparsity and credit assignment challenges in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1 points, while outperforming all baselines on group-level conversion-rate error and semantic trajectory distance and generalizing to out-of-distribution scenarios. In human Turing tests, annotators identified TRACER conversations at near-chance accuracy. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion and response quality via simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.