arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09764cs.CL

SocialRL:通过多轮强化学习与奖励设计提升大语言模型的社交智能

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

Jianing Wang, Xintao Wang, Aili Chen, Jie Shi, Hongcheng Guo, Jun Gao, Wenxuan Zhao, Chengkun Lang, Yuanli Guo, Yanghua Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

SocialRL提出多轮强化学习框架,通过PPO传播延迟奖励和六维过程奖励设计,解决现有方法短视问题,在社交对话基准上目标达成率平均提升9.2个百分点。

中文摘要 AI 辅助

社交智能使智能体能够解读社交情境、推断意图,并在持续对话中做出适应。随着语言模型成为自主协作者,社交智能对于构建有效且可信的人机交互至关重要。现有的强化学习方法优化单轮话语和稀疏的结果奖励,产生短视的策略,难以在多轮交互中管理目标与关系之间的张力。我们提出SocialRL,一个多轮强化学习框架,以解决这两个挑战。首先,我们应用基于PPO的多轮强化学习,将延迟的结果奖励传播回每一轮,从而实现长时程规划。其次,我们设计了六个过程奖励维度,捕捉目标与关系的权衡,包括目标推进、关系调谐、上下文连贯性等。一个奖励模型动态地为每个维度生成细粒度的评分标准,而一个阶段感知的权重调度在早期轮次优先关系建立,中期优先目标推进,后期优先平衡收尾。在多个社交对话基准上,SocialRL相比对应的基础模型,目标达成率平均提高了9.2个百分点。这些结果证明了SocialRL在合成和真实社交场景,以及标准和挑战性社交情境中的有效性。

英文摘要

Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to building effective and trustworthy human-AI interaction. Existing reinforcement learning methods optimize single-turn utterances and sparse outcome rewards, producing short-sighted policies that struggle to manage goal-relationship tensions across multi-turn interactions. We propose SocialRL, a multi-turn reinforcement learning framework addressing both challenges. First, we apply multi-turn reinforcement learning using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning. Second, we design six process reward dimensions capturing the goal-relationship trade-off, including goal advancement, relational attunement, contextual coherence, etc. A reward model dynamically generates fine-grained scoring criteria for each dimension, while a stage-aware weight schedule prioritizes relationship-building in early turns, goal advancement mid-way, and balanced closure late. Across multiple social-dialogue benchmarks, SocialRL improves Goal Achievement by an average of 9.2 percentage points over the corresponding Base models. These results demonstrate the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.

发表机构

  • Fudan University(复旦大学)
  • Hello Group(陌陌科技)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑