发表机构
Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SocialRL方法训练4B语言模型的社会推理能力,经多领域实验,其跨领域整合后模型平均效用达0.627,优于多数GPT系列模型,心智理论蒸馏可提升效用与泛化性。
AI 中文摘要
AI智能体越来越多地代表用户执行任务,例如安排会议、比较报价和讨价还价。这些由委托方驱动的任务通常会让智能体面临目标可能与委托方冲突的对方(另一个用户的智能体、卖家、招聘人员)。然而,让助手讨人喜欢的特质可能会使其成为糟糕的代理:一个友好、乐于助人的前沿模型可能会主动披露委托方的私人信息,并在首次遇到阻力时就让步。我们提出了SocialRL,一种直接训练社会推理的通用方法,并将其应用于一个4B模型,覆盖六个领域:Deal-or-No-Deal、CaSiNo、Craigslist、Job Interview、Calendar和Marketplace。每个领域都在相同方法下进行领域内训练,且每个策略都在全部六个领域接受评估。我们发现:(1)领域内训练达到了前沿水平:在保留场景中,该4B模型在每个领域的表现与GPT-5系列相当或更好,在谈判游戏中缩小了73%-122%的基线到前沿差距,其中78%的买方开局报价锚定在目标以下,而未训练模型的这一比例仅为3%;(2)跨领域迁移遵循游戏结构:结构配对的游戏会相互促进,广泛的多问题源会提升几乎所有领域,而结构孤立的游戏则无任何迁移;(3)基于该迁移结构,两种策略——级联RL和多教师在线策略蒸馏(OPD)——将各领域的专用模型整合为一个统一的4B模型,在全部六个环境中达到0.627的平均效用,与GPT-4.1(0.625)、GPT-5.1(0.619)和GPT-5.2(0.613)相当或更好;(4)明确的心智理论(ToM)支架仅在训练中发挥作用:蒸馏ToM轨迹而非仅蒸馏动作,可提升每个环境的效用并实现更好的跨领域泛化,且在两种ToM技能中,仅下一个动作预测能预测谈判结果。
英文摘要
AI agents increasingly act on their users' behalf, handling tasks such as scheduling meetings, comparing offers, and haggling over prices. These principal-driven tasks routinely place the agent across from a counterpart (another user's agent, a seller, a recruiter) whose goals may conflict with its principal's. Yet the dispositions that make an assistant pleasant can make it a poor delegate: a friendly, helpful frontier model may disclose its principal's private information unprompted and concede at the first sign of resistance. We present SocialRL, a general recipe that trains social reasoning directly, and apply it to a 4B model across six domains: Deal-or-No-Deal, CaSiNo, Craigslist, Job Interview, Calendar, and Marketplace. Every domain is trained in-domain under the same recipe, and every policy is evaluated on all six. We find that (1) in-domain training reaches the frontier: on held-out scenarios the 4B matches or exceeds the GPT-5 family per domain, closing 73-122% of the baseline-to-frontier gap on the negotiation games, with 78% of buyer openings anchoring below target versus 3% untrained; (2) cross-domain transfer follows game structure: structurally paired games lift each other, a broad multi-issue donor lifts nearly all domains, and structurally isolated games transfer nothing; (3) guided by this transfer structure, two strategies, cascade RL and multi-teacher on-policy distillation (OPD), consolidate the per-domain specialists into a single unified 4B that reaches 0.627 average utility across all six environments, matching or exceeding GPT-4.1 (0.625), GPT-5.1 (0.619), and GPT-5.2 (0.613); (4) an explicit theory-of-mind scaffold helps only through training: distilling the ToM trace, rather than actions alone, lifts utility on every environment and generalizes better across them, and of the two ToM skills, only next-action prediction predicts negotiation outcomes.
Comments25 pages, 3 figures