arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34249cs.AI

共情强化学习中的支持优先级演化

Evolving Support Priorities in Empathetic Reinforcement Learning

Pengyu Huang, Zhiyuan Han, Wenwen Tong, Hewei Guo, Jiangnan Chen, Sirui Chen, Lewei Lu, Beier Zhu, Xun Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对共情强化学习中支持优先级随对话状态演化而现有方法固定奖励的问题,提出CARE框架,通过上下文自适应评分标准动态调整奖励,在多个基准上显著提升性能。

中文摘要 AI 辅助

我们识别出共情强化学习中的一个根本性不匹配:支持优先级随对话状态演化,而现有方法通常优化在轮次间保持固定的预定义奖励规范。为建模这些演化的支持优先级,我们沿认知、情感和主动共情三个维度组织共情支持,并提出上下文自适应评分标准演化(CARE)。在每一轮,CARE通过调整这三个共情维度的权重及其细粒度评估标准来生成上下文自适应评分标准。评分标准生成器通过监督微调及随后的基于偏好的强化学习,利用轮级评分标准监督和人类偏好数据进行训练,然后作为在线共情强化学习的自适应奖励接口。集成到RLVER和MICA后,CARE在SentientBench、EQBench3和EMPA上,在三个独立的大语言模型评判下均达到最先进性能。值得注意的是,在EMPA上,CARE在所有三个评判下将EPM-Idx相对于最强基线至少提升13个点,包括在Gemini-2.5-Pro下从28.11提升至83.54。进一步分析表明,学习到的评分标准优先级在对话阶段和用户情绪间系统性地变化,证明CARE适应了随支持需求演化而被奖励的内容。

英文摘要

We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • SenseTime Research(商汤科技研究院)
  • Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

↑