利用层级推理和话语级目标奖励增强大型语言模型的社交智能
Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding
浏览论文内容
中文总结 AI 辅助
该研究针对LLMs社交互动不足的问题,提出TSR框架与LHRL-VGR算法,微调Qwen2.5-7B智能体后在SOTOPIA基准上超过GPT-4o基线7.32%,实现多智能体社交谈判的最优性能。
中文摘要 AI 辅助
大型语言模型(LLMs)在结构化任务中表现出色,但在动态社交互动中存在不足,这类互动的成功需要长期目标协调与快速适应。现有方法常对每轮话语应用统一的基于目标的奖励,忽视了每轮对话目标的特异性,也未考虑潜在策略的合理性。受计划行为理论启发,我们提出Think-Strategy-Response(TSR)框架,将社交对话分解为高层战略规划与低层语言执行两个层级阶段。为优化TSR,我们引入Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards(LHRL-VGR),这是一种新型算法,可根据目标达成分数的方差动态分配奖励,平衡目标完成度与策略一致性。在SOTOPIA基准测试上的实验表明,我们的方法对Qwen2.5-7B智能体进行微调后,目标完成成功率超过GPT-4o基线7.32%,在多智能体社交谈判任务中展现出最先进的性能。
英文摘要
Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and failing to account for the rationale of potential strategies. Inspired by the Theory of Planned Behavior, we propose the Think-Strategy-Response (TSR) framework, which decomposes social dialogue into two hierarchical stages: high-level strategic planning and low-level linguistic execution. To optimize TSR, we introduce Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR), a novel algorithm that dynamically routes rewards - balancing goal completion and strategy adherence - based on the variance of goal achievement scores. Experiments on the SOTOPIA benchmark show that our approach fine-tunes a Qwen2.5-7B agent to surpass the GPT-4o baseline by 7.32% in goal completion success, demonstrating state-of-the-art performance in multi-agent social negotiation tasks.
发表机构
- Alibaba Group(阿里巴巴集团)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。