AI 中文总结
本研究提出PsyAgentBench基准,通过因子设计区分LLM智能体对心理效应的识别与真实偏差,发现效应源于多种途径而非单一易感性,并建议用复制概况替代标量偏差分数。
AI 中文摘要
LLM产生与人类心理效应相关的反应模式,并不等同于该LLM具有这种偏差。我们提出了PsyAgentBench,一个在因子设计下对LLM智能体重新运行经典心理学实验的基准,该设计旨在区分这两者:每个范式在提示中明确标注范式名称(命名条件)或作为常规任务呈现(盲条件),并在任务的教科书字面版本(规范条件)或结构匹配的变体(反事实条件)上运行,该变体旨在减少与可能训练数据的词汇和场景重叠,并与角色操纵交叉。在五个已完成的范式中,在多达三个开放权重模型家族上评估,发布了41,904次试验,表面上类似人类的效应通过定性不同的途径产生,而非单一的易感性:范式标签门控与显式覆盖(阿希从众,gpt-oss-120B上从盲条件的0%到命名条件的83.3%),知识依赖的信号依赖(锚定,在基于事实的数值上恰好为零,而在虚构数量上几乎完全依赖,这一模式同样与理性使用唯一可用信号一致),在标签下对新内容的放大(框架效应),稳健的缺失(沉没成本),以及安全介导的选择,其中拒绝本身是主要发现(最小群体分配)。一句话的角色改变(宜人性,作为指令而非验证的特质操纵)根据效应的不同而消除、减弱或逆转这些效应,这反对任何单一反应偏差的解释。我们进一步形式化了心理学范式无法移植到LLM智能体的三种方式,并在两种情况下进行了实证记录:角色主导、群体崩溃和安全选择。我们认为标量偏差易感性分数掩盖了这种结构,并转而报告复制概况。
英文摘要
An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation. Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility: paradigm-label gating with explicit override (Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B), knowledge-dependent signal reliance (anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal), amplification on novel content under labeling (framing), robust absence (sunk cost), and safety-mediated selection where refusal itself is the primary finding (minimal-group allocation). A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.
Comments13 pages, 1 figure, 8 tables