arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34114cs.LGcond-mat.dis-nn

多目标强化学习框架中公平性的演化

Evolution of fairness in multi-objective reinforcement learning framework

Jingyi Zhang, Xin Ou, Guozhong Zheng, Shengfeng Deng, Jiqiang Zhang, Li Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出多目标强化学习框架,通过公平压力系数权衡收益与公平,模拟最后通牒博弈,发现中等压力下响应者变得宽容,揭示了策略逆转机制,扩展了RL范式以解释人类社交行为。

中文摘要 AI 辅助

公平作为一种基本的社会规范,其起源问题一直是一个长期存在的谜题。传统的博弈论模型在很大程度上依赖于“经济人”假设,即个体是纯粹理性和自利的,其行为仅仅是为了最大化物质收益。然而,这类解释忽视了人类决策的多维性质,人类的决策往往还受到经济激励之外的其他考量因素的影响。为了弥补这一不足,我们提出了一个多目标强化学习框架,将公平性的演化建模为物质收益最大化与公平驱动的道德行为之间的动态权衡,并通过一个公平压力系数进行调节。利用双目标Q学习最后通牒博弈的模拟,我们发现,公平压力的增加会促进公平结果的出现,这与预期一致。然而,引人注目的是,在中等压力下,响应者的行为发生了逆转:响应者变得“宽容”,接受低报价——这一模式与我们的日常经验相符。微观分析揭示,这种策略逆转源于收益最大化偏好与公平导向偏好之间的竞争。我们进一步将框架扩展到不对称设置,其中提议者和响应者赋予两个目标不同的权重。总体而言,我们的工作将强化学习范式从单目标扩展到多目标形式,为阐明更广泛的人类社会行为提供了一种多功能工具。

英文摘要

Fairness, as a fundamental social norm, continues to pose a longstanding puzzle regarding its emergence. Traditional game-theoretic models largely rely on the assumption of \emph{Homo economicus}, wherein individuals are purely rational and self-interested, acting solely to maximize material payoffs. Such accounts, however, overlook the multidimensional nature of human decision-making, which is often shaped also by other considerations beyond economic incentives. To address this gap, we propose a multi-objective reinforcement learning framework that models the evolution of fairness as a dynamic trade-off between material payoff maximization and fairness-driven moral behavior, regulated by a fairness pressure coefficient. Using simulations of a two-objective Q-learning ultimatum game, we find that increased fairness pressure promotes fair outcomes, as expected. Strikingly, however, under moderate pressure, responder behavior reverses: responders become ``forgiving" by accepting low offers -- a pattern in line with our daily experience. Microscopic analyses reveal that this strategy reversal stems from competition between payoff-maximizing and fairness-oriented preferences. We further extend our framework to an asymmetric setting, where proposers and responders assign different weights to the two objectives. Overall, our work expands the reinforcement learning paradigm from a single-objective to a multi-objective formulation, offering a versatile tool for elucidating a broader range of human social behaviors.

发表机构

  • Shaanxi Normal University(陕西师范大学)
  • Inner Mongolia University(内蒙古大学)
  • Ningxia University(宁夏大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑