arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于认知状态转移的主动对话策略优化

Proactive Dialogue Policy Optimization via Cognitive-State Transition

Minghui Ma, Mengqi Chen, Bin Guo, Jingqi Liu

arXiv 2609.34948首次发表:更新:

发表机构

Northwestern Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对主动对话中用户认知演化建模不足及策略标签粒度粗糙的问题,联合设计认知用户模拟器Cog-Sim与认知状态转移驱动的策略优化CSTPO,通过层次化策略-话语表示和双级优势估计,显著提升策略优化效果,使Qwen3-14B达到GPT-5.5规划方法的水平。

AI 中文摘要

主动对话要求智能体在多个轮次中朝着任务目标推进的同时,持续根据用户反馈调整其策略。为了超越在静态数据集上的模仿学习,近期方法使用用户模拟器收集交互数据进行策略优化。然而,许多模拟器并未显式建模用户认知的演化,限制了跨轮次反馈的一致性和状态依赖性。此外,仅用高层策略标签表示每个动作会忽略庞大的话语空间,无法区分同一策略的不同实现。为此,我们联合设计了认知用户模拟器(Cog-Sim)和认知状态转移驱动的策略优化(CSTPO)。Cog-Sim维护用户的认知和情感状态,并通过跨轮次的受约束状态转移生成响应,使反馈既依赖于实际话语,也依赖于用户当前状态。CSTPO将每个动作组织为层次化的策略-话语表示:高层策略标签约束话语采样,话语在每个标签内进行优化。稀疏完整分支采样复用共享对话前缀,并分别估计策略级和话语级优势,从而在两个层面实现细粒度优化。在三个任务中,Cog-Sim表现出单调的剂量-反应关系,并且在自然度上优于基于提示的模拟器。CSTPO将Qwen3-14B的性能提升至与基于GPT-5.5的规划方法相当的水平。

英文摘要

Proactive dialogue requires agents to continually adapt their policies to user feedback while progressing toward task objectives over multiple turns. To move beyond imitation learning on static datasets, recent approaches use user simulators to collect interactive data for policy optimization. However, many simulators do not explicitly model the evolution of user cognition, limiting the consistency and state dependence of feedback across turns. Moreover, representing each action only by a high-level strategy label overlooks the large utterance space and cannot distinguish alternative realizations of the same strategy. To this end, we jointly design a $\textbf{Cog}$nitive User $\textbf{Sim}$ulator $\textbf{(Cog-Sim)}$ and $\textbf{C}$ognitive-$\textbf{S}$tate $\textbf{T}$ransition--Driven $\textbf{P}$olicy $\textbf{O}$ptimization $\textbf{(CSTPO)}$. Cog-Sim maintains the user's cognitive and affective states and generates responses through constrained state transitions across turns, so feedback depends on both the realized utterance and the user's current state. CSTPO organizes each action as a hierarchical strategy--utterance representation: a high-level strategy label constrains utterance sampling, and utterances are optimized within each label. Sparse complete-branch sampling reuses shared dialogue prefixes and estimates separate strategy-level and utterance-level advantages, enabling fine-grained optimization at both levels. Across three tasks, Cog-Sim exhibits monotonic dose--response relationships and is preferred over prompt-based simulators for naturalness. CSTPO improves Qwen3-14B's performance to a level comparable to that of GPT-5.5-based planning methods.

Comments30pages, 9figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑