arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉语言模型中的奖励估值:快感缺乏背后的因果机制

Reward Valuation in Large Language Models: Causal Induction of Anhedonia

Melika Honarmand, Samin Mahdipour Aghabagher, Martin Schrimpf

arXiv 2607.06626首次发表:更新:

发表机构

EPFL(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨视觉语言模型中奖励估值情况,基于神经科学观点,通过有针对性扰动识别奖励预期单元,测试其因果作用,发现模型存在奖励估值和预期缺陷,结果反映了人工智能模型与人类平行的奖励估值回路。

AI 中文摘要

近期视觉语言模型捕捉到人类认知中日益复杂的方面。本文探讨这种一致性是否延伸到奖励估值,在基于评估重度抑郁症中快感缺乏和动机缺陷的临床试验构建的机制框架中进行评估。在大脑中,快感缺乏常与伏隔核(NAc)及更广泛的多巴胺能奖励系统失调相关。虽神经成像已定位这些缺陷,但建立NAc活动与特定行为症状之间的因果联系仍是挑战。我们利用神经科学的这些观点在视觉语言模型中功能识别奖励预期单元,并通过有针对性的扰动测试其因果作用。扰动NAc选择性单元会引发类似人类快感缺乏的行为效应:模型在基于努力的决策任务中转向低努力、低奖励选项。关键的是,我们的结果反映了奖励估值和预期的特定缺陷而非任务能力丧失:当基于奖励的选择被移除时,受扰动模型保持基线表现。这种诱发的脆弱性进一步与包括DARS和MAP - SR在内的临床快感缺乏和动机量表一致。这些结果揭示了人工智能模型中与人类平行的奖励估值回路。

英文摘要

Recent frontier models mimic complex aspects of human cognition. Here we ask whether this alignment extends into reward valuation, which we assess in a mechanistic framework. Specifically, we use clinical tests that were developed to evaluate anhedonia in human subjects with major depressive disorders. Mechanistically, anhedonia is frequently associated with dysregulation in the Nucleus Accumbens (NAc) and the broader dopaminergic reward system. While neuroimaging has localized these deficits, establishing a causal link between NAc activity and specific behavioral symptoms remains a challenge. We use these ideas from neuroscience to functionally identify reward-anticipatory units in state-of-the-art AI models, and evaluate their causal involvement via targeted perturbations. We find that not only are such model units predictive of NAc brain recordings, their perturbation also induces behavioral effects mirroring human anhedonia: the model opts for low-effort, low-reward tasks in effort-based decision-making paradigms. Crucially, our results demonstrate that this represents a specific deficit in self-centered reward valuation and anticipation--rather than a loss of task capability, reward calculation, or effort avoidance. This induced vulnerability aligns with clinical measures of anhedonia and motivation in humans, such as DARS and MAP-SR, instruments that contain no reward-related vocabulary, ruling out a purely lexical account of the perturbation effect. Taken together, our results suggest reward valuation circuits in AI models that functionally mimic those in humans.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑