发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM社交适应性训练中真实交互数据稀缺和偏好潜在的问题,提出基于人格驱动模拟环境的偏好批处理GRPO算法,利用跨用户桶归一化优势稳定训练,实证优于强RL基线。
AI 中文摘要
构建不仅在行为上正确,而且在社交上表现良好的LLM,需要的不仅仅是产生局部有帮助的回应。一个具有社交能力的智能体必须推断用户未明说的目标,尊重他们的偏好,并随着对话的展开而适应。这些行为本质上是多轮且社交性的,使得它们难以优化:真实的交互数据稀缺,且用户偏好通常是潜在的而非直接可观察的。为了应对这些挑战,我们基于一个人格驱动的社会模拟环境(包括一个人格库、基于LLM的用户模拟器,以及一个范围在[0, 1]的用户满意度评分系统),引入了偏好批处理GRPO(PB-GRPO),一种从对话级反馈中学习社会适应性策略的后训练算法。与原始GRPO相比,PB-GRPO使用跨具有相似偏好的用户桶估计的归一化来计算优势,从而在多样化的社会群体中稳定训练。实证证据表明,在我们的模拟环境中,PB-GRPO在强强化学习基线上改善了模型的社会行为。
英文摘要
Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently multi-turn and social, making them hard to optimize: real interaction data is scarce, and user preferences are typically latent rather than directly observable. To address these challenges, we build on a persona-driven social simulation environment (consisting of a persona library, LLM-based user simulators, and a user-satisfaction scoring system ranging from [0, 1]), to introduce preference-batched GRPO (PB-GRPO), a post-training algorithm that learns socially adaptive policies from conversation-level feedback. Compared to vanilla GRPO, PB-GRPO computes advantages using a normalization estimated across a bucket of users with similar preferences, stabilizing training across a diverse social population. Empirical evidence shows that PB-GRPO improves models' social behavior over strong reinforcement learning baselines in our simulated environment.