AI 中文总结
研究在强化学习智能体中建模心理障碍,将其重新表述为对认知评估信号的剂量可控操纵,通过多个旋钮表示多种障碍,呈现分级剂量反应,还发现障碍自组织成二维空间、旋钮对不同障碍有不同影响及旋钮相互作用可预测共病等,且框架具有通用性。
AI 中文摘要
在人工智能体中建模心理障碍,既为计算精神病学提供了测试平台,也为情感控制的失败模式提供了视角。以往工作通过手工调整奖励塑造在强化学习智能体中诱导出一两种障碍,事后标记行为并报告单次运行结果。我们将障碍建模重新表述为在评估引导的近端策略优化(PPO)智能体中对认知评估信号进行剂量可控的操纵,将七种障碍(焦虑、躁狂、强迫检查、抑郁、冲动、成瘾和创伤后应激障碍)各自表示为基于计算精神病学解释的单个旋钮,每种症状通过预先注册的测定法进行测量并映射到公认的范式。在超过一千次运行(10个种子、4个对照、95%置信区间)中,每种障碍都呈现出分级、单调的剂量反应,没有对照能重现这种反应。除了这些诱导效应外,还出现了三个未写入奖励中的发现:这些障碍自组织成一个二维情感空间,其中躁狂与焦虑镜像;移除一个旋钮可缓解奖励扭曲障碍(躁狂、检查、成瘾),但不能缓解回避障碍(焦虑、创伤后应激障碍),后者在分级暴露课程下反而恢复;两个同时的旋钮相互作用是非加性的,产生可测试的共病预测。评估权重因此参数化了一个可控的情感表型空间,其中诱导障碍的相同旋钮可以模拟其治疗。我们还表明,三个障碍旋钮(抑郁、成瘾、焦虑)可转移到具有标准卷积智能体且没有评估评论家的三维像素环境(MiniWorld)中,跨域确认了跨测定解离,表明该框架不限于网格世界或PPO的评估评论家。
英文摘要
Modelling psychological disorders in artificial agents offers a testbed for computational psychiatry and a lens on affective-control failure modes. Prior work induces one or two disorders by hand-tuned reward shaping, labels the behaviour post hoc, and reports single runs. We recast disorder modelling as dose-controllable manipulation of cognitive appraisal signals in an appraisal-guided PPO agent, expressing seven disorders (anxiety, mania, obsessive-compulsive checking, depression, impulsivity, addiction, and post-traumatic stress) each as a single knob grounded in a computational psychiatry account, with each symptom measured by a preregistered assay. Across more than a thousand runs (10 seeds, four controls, 95% confidence intervals) every disorder shows a graded, monotone dose-response that no control reproduces. Beyond these induced effects, three findings emerge that were not written into the reward: disorders self-organise into a two-dimensional affective space in which mania mirrors anxiety; removing a knob remits reward-distortion disorders (mania, checking, addiction) but not avoidance disorders (anxiety, PTSD), which recover under a graded exposure curriculum; and two simultaneous knobs interact nonadditively, yielding testable comorbidity predictions. The depression and addiction knobs further reproduce their double dissociation in a 3D pixel environment (MiniWorld) with a standard convolutional agent and no appraisal critic, showing the framework generalises beyond grid worlds.
Comments15 pages, 8 figures, 6 tables