arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07881cs.AIcs.MA

自参照社会偏好:无需观察他人奖励的合作

Self-Referenced Social Preferences: Cooperation without Observing Others Rewards

Mohamed Ayman Mohamed, Harshil Kotamreddy, Marcos Menon Jose

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出自参照社会偏好,使智能体无需观察他人奖励,通过自身奖励模型评估他人结果以促进多智能体合作,并在三个连续社会困境中验证了其有效性。

中文摘要 AI 辅助

社会偏好可以促进多智能体强化学习中的合作,但现有方法通常要求智能体观察同伴的奖励。然而,在许多现实世界的交互中,智能体可以像人类一样,观察他人的行为和结果,而无需获取其私有奖励信号。我们引入了自参照社会偏好,其中每个智能体学习自身奖励的模型,将其应用于其他智能体观察到的转移,以从自身视角评估他们的结果,并将这些自参照评估输入到标准社会偏好中。我们研究了整合这些评估的两种方式:修改学习奖励,或使用它们来加权策略更新。我们在三个连续社会困境中评估了该方法:Escape Room、Clean Up 和 Commons Harvest,这些困境分别要求自愿参与、公共物品贡献和资源节制。在所有三个环境中,智能体无需观察他人奖励即可学习合作行为,包括在独立学习者无法合作的环境中,并且与能够获取真实奖励的智能体相比,通常能实现更公平的共同产出分配。有效的整合点取决于社会偏好:不公平厌恶在奖励中与价值前瞻结合效果最佳,而纯粹利他偏好则受益于策略更新加权。在部分可观测性下,策略更新方法继续支持合作。这些结果表明,显式访问其他智能体的奖励信号并非学习合作行为的必要条件:社会偏好可以基于从他人观察行为中得出的自参照结果评估来建立。

英文摘要

Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approach on three sequential social dilemmas, Escape Room, Clean Up, and Commons Harvest, which require volunteering, public-good contribution, and resource restraint, respectively. Across all three environments, agents learn cooperative behavior without observing others' rewards, including in settings where independent learners fail to cooperate, and frequently achieve more equitable divisions of jointly produced returns than agents with access to true rewards. The effective integration point depends on the social preference: inequity aversion works best in the reward together with a value look-ahead, whereas a purely benevolent preference benefits from policy-update weighting. Under partial observability, the policy-update approach continues to support cooperation. These results show that explicit access to other agents' reward signals is not necessary for learning cooperative behavior: social preferences can instead be grounded in self-referenced assessments of others' outcomes derived from their observed behavior.

发表机构

  • Amazon(亚马逊)
  • Nvidia(英伟达)
  • Itaú Unibanco(伊塔乌联合银行)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑