AI 中文总结
该研究针对大语言模型共情能力不足问题,提出双循环自演化框架,在SAGE数据集上显著提升Qwen3-8B的共情对话性能,优于对比方法。
AI 中文摘要
大语言模型已展现出对话能力,但共情能力仍是挑战。共情支持本质上是多轮且路径依赖的:用户逐步披露顾虑、情感随时间演变,早期响应会影响信任与接受度。带有可验证情感奖励的强化学习为长程交互提供了可扩展监督,但现有方法在演化对话策略时保持训练交互分布固定,导致策略能力与训练经验不匹配。本文提出由可验证情感反馈驱动的双循环自演化框架:在用户模拟器和验证器冻结的情况下,内循环用连续情感奖励优化多轮策略,外循环利用相同结果估计策略相关的交互效用并适配经验;为从稀疏随机rollout中获取估计,框架在每组内保持场景和交互状态恒定,优先选择组通过率接近策略能力边界的条件,分层控制器在支持意图间共享证据,不确定性引导探索和均匀回放防止过早排除。该框架生成轨迹并闭合双循环,未增加rollout预算;在SAGE数据集上,其将Qwen3-8B的Overall指标从53.87提升至79.24,比协议匹配的均匀情感奖励强化学习高出7.23个百分点。
英文摘要
Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently multi-turn and path-dependent: users disclose concerns gradually, emotions evolve over time, and early responses shape trust and receptivity. Reinforcement learning with verifiable emotion rewards provides scalable supervision for long-horizon interactions. However, existing methods evolve the dialogue policy while keeping its training interaction distribution fixed, creating a mismatch between policy competence and training experience. We introduce a dual-loop self-evolution framework driven by verifiable emotion feedback. With the user simulator and verifier frozen, the inner loop optimizes the multi-turn policy using continuous emotion rewards, while the outer loop uses the same outcomes to estimate policy-relative interaction utility and adapt experience. To obtain estimates from sparse, stochastic rollouts, the framework holds the scenario and interaction state constant within each group and prioritizes conditions whose group pass rates lie near the policy's competence boundary. A hierarchical controller shares evidence across support intents, while uncertainty-guided exploration and uniform rehearsal prevent premature exclusion. The resulting distribution generates trajectories, closing both loops without increasing the rollout budget. On SAGE, our framework raises Qwen3-8B Overall from 53.87 to 79.24 and outperforms protocol-matched uniform emotion-reward reinforcement learning by 7.23 points.
Comments10 pages, 4 figures, 6 tables