arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于比较证据的社交智能强化学习

Reinforcement Learning with Comparative Evidence for Social Intelligence

Keane Ong, Yuriel Ryan, Sabri Boughorbel, Vladimir Necula, Jack Wei Lun Shi, Rui Mao, Roy Ka-Wei Lee, Adriel Kuek, Nancy F. Chen, Erik Cambria, Gianmarco Mengaldo, Paul Pu Liang

arXiv 2610.04072首次发表:更新:

发表机构

MIT; NUS; Prince Sattam bin Abdulaziz University; University of Cambridge; NTU; UBC; DSO National Laboratories; A*STAR(麻省理工学院; 新加坡国立大学; 沙特国王大学; 剑桥大学; 南洋理工大学; 不列颠哥伦比亚大学; 国防科技局国家实验室; 新加坡科技研究局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对社交智能训练依赖人工标注且社会预测缺乏验证的问题,提出基于比较证据的强化学习(RLCE),通过构建和聚合证据测试从无标注数据中学习,在四个基准上超越七种无标签方法,最高提升18.93分。

AI 中文摘要

开发具有社交智能的人工智能仍然严重依赖人工标注的数据,这限制了模型能够获得的社会理解的规模和广度。从无标注数据中提取训练信号的方法提供了一条超越这种依赖的路径,但社会预测缺乏数学和编码领域中可用的验证预言机。此外,核心社会目标,如情感、意图、偏好和语用意义,往往是模糊的。同一行为可能支持多种合理的解释,这使得难以验证哪种解释最受支持。为了应对这一挑战,我们引入了基于比较证据的强化学习(RLCE),这是一种从无标注训练数据中学习社会理解的强化学习方法,无需从真实标注构建奖励。给定一个 rollout 组中的不同答案,RLCE 构建证据测试,识别出支持一个答案优于另一个答案的可观察证据,针对输入样本验证这些测试,并聚合测试结果以确定最受支持的解释。随着策略产生新的答案,测试会被重新生成,使其能够随策略演化。在涵盖情感、语用学、交际意图和偏好的四个基准测试中,RLCE 在七种不使用真实训练标签作为奖励的方法中取得了最强性能,这些方法包括共识、策略 LLM 法官验证、多模态协同进化和基于量规的奖励。相对于最强基线的提升最高达 +18.93 分。进一步的分析表明,与所比较的量规方法相比,RLCE 在正确和错误预测之间的奖励差异占比更大,能够推翻错误的策略衍生偏好,并受益于成对测试构建、组合测试聚合和在线策略测试演化。

英文摘要

Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but social predictions lack the verification oracles available in mathematics and coding. Moreover, core social targets such as affect, intent, preference, and pragmatic meaning are often ambiguous. The same behavior can support multiple plausible interpretations, making it difficult to verify which is best supported. To address this challenge, we introduce Reinforcement Learning with Comparative Evidence (RLCE), a reinforcement learning method that learns social understanding from unlabeled training data without constructing rewards from ground-truth annotations. Given distinct answers in a rollout group, RLCE constructs evidence tests that identify observable evidence favoring an answer over another, validates these tests against the input sample, and aggregates test outcomes to determine the best-supported interpretation. Tests are regenerated as the policy produces new answers, enabling them to evolve with the policy. Across four benchmarks spanning affect, pragmatics, communicative intent, and preference, RLCE attains the strongest performance among seven methods that use no ground-truth training labels for rewards, including consensus, policy LLM-judge verification, multimodal co-evolution, and rubric-based rewards. Gains over the strongest baseline reach up to +18.93 points. Analyses further show that RLCE exhibits a larger share of reward variation between correct and incorrect predictions than compared rubric methods, can overturn erroneous policy-derived preferences, and benefits from pairwise test construction, compositional test aggregation, and on-policy test evolution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑