发表机构
Geely AI Lab; Beijing Institute of Technology; Peking University(吉利AI实验室; 北京理工大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究受人类多层时间抽象启发,提出ToSCA框架,采用两级分层强化学习结合双粒度奖励机制,在两类对话任务中性能优于基线方法。
AI 中文摘要
人类在日常交互与思考中具有多层时间抽象,如概念感知与策略规划。受此启发,我们提出一种用于对话智能体的两级分层强化学习(RL)框架,以弥合此前基于token级或utterance级RL方法的差距。该框架基于两级MDP构建,token级响应解码以utterance级动作(显式文本策略)为条件。基于理论推导与效率考量,我们使用DQN求解高级评论者,使用PPO求解低级演员-评论者。为进一步缓解奖励稀疏性并促进收敛,我们还设计了双粒度奖励机制,将utterance级满意度评分与token级内在动机及K-L惩罚相结合。在日常对话与情感支持对话上的实验表明,我们的方法在策略确定与响应质量上优于多种基线方法。我们的实现可在该https URL获取。
英文摘要
Humans naturally exhibit multiple forms of abstraction in reasoning and interaction, including temporal abstraction across decision timescales and strategic abstraction over communicative intents. Inspired by these complementary abstractions, we propose a two-level hierarchical reinforcement learning (HRL) framework for conversational agents that bridges the gap between existing token-level and utterance-level RL methods. Built upon a two-level Markov decision process (MDP), our framework conditions token-level response generation on utterance-level actions represented by explicit textual strategies. Based on theoretical analysis and efficiency considerations, we employ DQN to optimize the high-level Q-network and PPO to train the low-level actor-critic. To further alleviate reward sparsity and facilitate convergence, we introduce a dual-granularity reward mechanism that combines the utterance-level satisfaction score with token-level intrinsic self-consistency and a KL-divergence penalty. Experiments on both daily-life and emotional support conversations demonstrate that our method consistently outperforms a wide range of baselines in both strategy determination and response quality. Our implementation is available at https://github.com/AaronJi/ToSCA.
CommentsAccepted by EMNLP 2026 Findings