arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

校准人工愧疚:用于亲社会多智能体强化学习的神经基础奖励塑造

Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning

Aaditya Mehta, Arya Shah

arXiv 2608.04663首次发表:更新:

发表机构

Mahatma Gandhi International School(圣雄甘地国际学校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究从人类神经行为数据校准愧疚信号作为亲社会奖励权重,嵌入多智能体强化学习环境后,其智能体的社会选择率最贴合人类数据,为亲社会奖励塑造提供了定量约束。

AI 中文摘要

合作型多智能体强化学习常向个体奖励中添加社会项,但这些项的规模通常由手动选择。我们探究是否可从人类神经与行为数据中校准愧疚信号,并将其迁移至人工智能体。使用公开的SoDec责任fMRI数据集(含40名参与者),我们拟合了瞬间快乐变化对结果类型计数的被试固定效应回归,恢复出愧疚权重为“伴侣负性”减去“社会负性”的对比值($\boldsymbol{\tilde{w}}=1.118$,Cohen's $d=0.214$)。我们将该权重嵌入双智能体“社会彩票”环境,在四种塑造机制下训练独立的近端策略优化(Proximal Policy Optimization)演员-评论家:神经校准型、均匀常数型、零(自私型)及单位系数神谕型。每种条件下各进行1000次评估回合,结果显示校准型智能体最贴合人类的社会安全选择率(0.459,人类为0.484;KL散度=0.0012),其余三种条件的KL散度偏差为1至3个数量级。因此,人类神经行为先验可作为亲社会奖励塑造的定量约束。

英文摘要

Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand. We ask whether a guilt signal can instead be calibrated from human neural and behavioural data and transferred to artificial agents. Using the public SoDec responsibility fMRI dataset (40 participants), we fit a subject-fixed-effects regression of momentary-happiness changes on outcome-type counts and recover a guilt weight as the Partner-negative minus Social-negative contrast ($\hat{w}=1.118$, Cohen's $d=0.214$). We embed this weight in a two-agent Social Lottery environment and train independent Proximal Policy Optimization actor-critics under four shaping regimes: neurally calibrated, uniform constant, zero (selfish), and a unit-coefficient oracle. Across 1{,}000 evaluation episodes per condition, the calibrated agents track the human Social safe-choice rate most closely ($0.459$ vs.\ human $0.484$; $\mathrm{KL}=0.0012$), while the other three conditions deviate by one to three orders of magnitude in KL. Human neurobehavioural priors can therefore act as quantitative constraints on prosocial reward shaping.

Comments12 pages, 6 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑