arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21925cs.AI

ESCRAG-R1:用于情感支持对话的检索增强强化学习

ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation

Weichu Liu, Yuxuan Hu, Yirong Sun, Ningning Mao, Ziyun Zhang, Jian Chen, Mingyang Xu, Qishan Zhong, Chengming Li

首次发表
浏览论文内容

中文总结 AI 辅助

ESCRAG-R1是整合检索式心理指导与GRPO的统一框架,构建了ESC-Preference数据集,在情感支持对话任务中显著优于现有基线。

中文摘要 AI 辅助

情感支持对话(ESC)系统旨在通过平衡专业治疗能力与自然共情来提供全面支持。然而,现有方法难以同时实现结构化、阶段感知的推理以及共情与专业知识的无缝对齐,常导致临床策略与通用安慰的人为拼接。为克服这些局限,我们提出ESCRAG-R1,这是一个将基于检索的心理指导整合到分组相对策略优化(GRPO)中的统一框架。通过在强化学习循环中引入检索,ESCRAG-R1将外部知识转化为鲁棒的学习信号,在生成前激发明确的内部推理,并从根本上重塑模型的内部策略。为提供该优化所需的可靠监督,我们构建了ESC-Preference,这是一个基于来访者-咨询师-评估者(Client--Counselor--Judge)评估框架的高质量数据集,可提供精确、感知共情的奖励信号。大量实验表明,ESCRAG-R1通过减少表面拼接并实现专业指导与共情表达的自然整合,显著优于现有基线。代码与数据集已发布于此https URL。

英文摘要

Emotional Support Conversation (ESC) systems aim to provide holistic support by balancing professional therapeutic competence with natural empathy. However, existing methods struggle to simultaneously achieve structured, stage-aware reasoning and seamless empathy-expertise alignment, often resulting in an artificial splicing of clinical strategies and generic reassurance. To overcome these limitations, we propose ESCRAG-R1, a unified framework that integrates retrieval-based psychological guidance into Group Relative Policy Optimization (GRPO). By incorporating retrieval into the reinforcement learning loop, ESCRAG-R1 transforms external knowledge into a robust learning signal that stimulates explicit internal reasoning prior to generation and fundamentally reshapes the model's internal policy. To provide the reliable supervision required for this optimization, we construct ESC-Preference, a high-quality dataset based on a Client--Counselor--Judge evaluation framework that delivers precise, empathy-aware reward signals. Extensive experiments demonstrate that ESCRAG-R1 significantly outperforms existing baselines by mitigating superficial splicing and realizing a natural integration of professional guidance and empathetic expression. Code and datasets are released at https://github.com/Matcha-Liu/ESCRAG-R1.

发表机构

  • Beijing Institute of Technology(北京理工大学)
  • Shenzhen MSU-BIT University(深圳北理莫斯科大学)
  • City University of Hong Kong(香港城市大学)
  • Shenzhen University of Advanced Technology(深圳理工大学)
  • Beijing Normal University(北京师范大学)
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑