arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为何要研究涌现行为?你可以对其进行调控!让多智能体系统与奖励预测对齐

Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction

Assaf Caftory, Almog Zemach, Moshe Butman, Doron Friedman

arXiv 2608.07280首次发表:更新:

AI 中文总结

本文提出多智能体奖励预测(MARP)框架,通过学习共享奖励模型调控多智能体涌现行为,在收获游戏中验证其能使智能体行为对齐多样社会目标,为调控涌现行为提供了数据驱动的方法。

AI 中文摘要

多智能体模拟被广泛用于研究复杂的社会和生态系统,其中丰富且往往出人意料的涌现行为源于局部交互。大量前期工作聚焦于分析不同领域的这类涌现动态。在本文中,我们不再局限于分析涌现行为,而是引入一种基于学习的机制,通过社会奖励建模主动塑造涌现行为。我们提出了多智能体奖励预测(Multi-Agent Reward Prediction, MARP),这是一种将基于偏好的奖励建模扩展至多智能体强化学习的简单框架。尽管该框架旨在适用于各类多智能体场景,但本次实证验证仅局限于单一环境,因此我们将MARP作为所研究领域内的概念验证。MARP不依赖手工设计的奖励,而是从集体结果的回合级评估中学习共享奖励模型,使去中心化智能体能够将自身行为与全局社会目标对齐。我们在收获游戏(Harvest Game)中研究了MARP,这是一种建模公共池资源管理及相关现实挑战的典型序列社会困境。结果表明,与基于标准奖励的基线方法相比,MARP可被调整为产生与目标社会指标更紧密对齐的行为,同时学习到的奖励模型无需显式编程即可捕捉微妙的环境结构。至关重要的是,MARP在单一训练机制内支持多个及复合社会目标。仅通过修改高级评估指标,同一框架即可无缝将智能体行为与多样化目标对齐,包括可持续性、平等性、和平,以及个体与群体层面目标的组合。这些发现表明,涌现多智能体行为不仅可作为一种待研究的现象,还可作为基于原则的数据驱动调控的目标。

英文摘要

Multi-agent simulations are widely used to study complex social and ecological systems, where rich and often unexpected emergent behaviors arise from local interactions. A large body of prior work has focused on analyzing such emergent dynamics across domains. In this paper, we move beyond analyzing emergent behavior and introduce a learning-based mechanism for actively shaping it via social reward modeling. We introduce Multi-Agent Reward Prediction (MARP), a simple framework that extends preference-based reward modeling to multi-agent reinforcement learning. While the framework is designed to be applicable across multi-agent settings, the present empirical validation is limited to a single environment, and we therefore present MARP as a proof of concept within the studied domain. Rather than relying on handcrafted rewards, MARP learns a shared reward model from episode-level evaluations of collective outcomes, enabling decentralized agents to align their behavior with global social objectives. We study MARP in the Harvest Game, a canonical sequential social dilemma modeling common-pool resource management and related real-world challenges. Our results show that MARP can be tuned to produce behavior that is more closely aligned with target social metrics than standard reward-based baselines, while the learned reward model captures subtle environmental structure without explicit programming. Crucially, MARP supports multiple and composite social objectives within a single training regime. By modifying only the high-level evaluation metric, the same framework seamlessly aligns agent behavior with diverse goals, including sustainability, equality, and peace, as well as combinations of individual and group-level objectives. These findings demonstrate that emergent multi-agent behavior can be treated not only as a phenomenon to study, but as a target of principled, data-driven regulation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑