arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

意图比回应更重要:超越回应模仿的可控用户模拟

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

Bo Wang, Ruixing Zhang, Yunqi Liu, Yang Zhang, Liangzhe Han, Tongyu Zhu, Leilei Sun

arXiv 2608.09420首次发表:更新:

发表机构

Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出UserIDA模型,将交互意图作为每轮明确指令,在LMSYS-USP上实现86.6%意图准确率,大幅优于基线,为用户模拟提供了与回应保真度互补的每轮意图控制维度。

AI 中文摘要

用户模拟器被广泛用作训练和评估交互式助手的可扩展环境。生成下一轮用户对话本质上是一对多的:相同的用户画像和对话上下文可能支持具有不同局部交互意图的多个合理延续。流畅的回应可能会通过不恰当的意图推进对话,比如接受而非修复。我们的核心见解是,可控用户模拟应将下一轮用户对话应实现的局部交互意图与该意图如何用语言表达分离开来。我们引入UserIDA(用户意图-指令对齐),它将交互意图作为每轮明确的指令。UserIDA定义了六向意图接口,通过监督微调学习指令条件生成,并在基于组的强化学习期间使用意图校准的策略优化。奖励机制保留复合回应质量,同时确保混合组中违反意图的候选排在符合意图的候选之后。在LMSYS-USP上,UserIDA实现了86.6%的意图准确率,比最强的专用用户模拟器基线高出24.3个百分点,同时提升了语义和风格相似性。在上下文内干预中,它在91.7%的评估对话状态中实现了至少六个目标意图中的四个,而最强的外部基线仅为22.9%。这些结果表明,每轮意图控制是用户模拟中与回应保真度互补的维度。

英文摘要

User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support multiple plausible continuations with different local interaction intents. A fluent response may therefore advance the dialogue through an inappropriate intent, such as acceptance rather than repair. Our key insight is that controllable user simulation should separate which local interaction intent the next user turn should realize from how that intent is expressed in language. We introduce UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive. UserIDA defines a six-way intent interface, learns directive-conditioned generation through supervised fine-tuning, and uses intent-calibrated policy optimization during group-based reinforcement learning. The reward preserves composite response quality while ensuring that intent-violating candidates rank below compliant alternatives in mixed groups. On LMSYS-USP, UserIDA achieves 86.6\% intent accuracy, outperforming the strongest dedicated user-simulator baseline by 24.3 percentage points while improving semantic and stylistic similarity. In within-context interventions, it realizes at least four of the six target intents in 91.7\% of evaluated dialogue states, compared with 22.9\% for the strongest external baseline. These results establish per-turn intent control as a complementary dimension to response fidelity in user simulation.

Comments26 pages, 7 figures, 16 tables. Code: https://github.com/ptwang773/UserIDA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑