arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从心智到模型:心理学与大语言模型(LLM)行为的交叉

From Minds to Models: The Intersection of Psychology and LLM Behaviours

Oliver Guidetti, Reza Ryan

arXiv 2607.27579首次发表:更新:

AI 中文总结

本研究结合心理学方法,测试ChatGPT等LLM的情感差异,发现种族条件的情感效应薄弱且依赖分析方法,提出需跨学科开发模型偏差的行为测量方法。

AI 中文摘要

大语言模型(LLM)因决策过程复杂、非线性且难以解释,常被拿来与人类心智作比较。用于研究不可观测心理过程的心理学方法,或可用于检验LLM行为,尤其在政府与医疗领域。本研究基于内隐联想测验(Implicit Association Test)的提示词适配方法,测试ChatGPT在开放式文本中是否会因种族条件产生情感差异。将14个基础问题与8个种族类别、1个无种族控制项交叉,生成126条提示词,每条分别提交给GPT-3.5T、GPT-4和GPT-4T各一次,共得到378条响应。情感得分由分类标签与来源得分推导:正面标签保留来源得分,负面标签取其负值,中性响应编码为0。双向方差分析(ANOVA)发现种族条件存在小的主效应,F(8, 351)=2.04,p=0.042,偏η²=0.044;但模型无主效应,F(2, 351)=0.07,p=0.933;种族条件与模型无交互效应,F(16, 351)=0.23,p=0.999。然而,该效应在秩转换敏感性分析中未被保留,F(8, 351)=1.53,p=0.145;经Tukey校正的比较未发现显著成对差异。未校正的欧洲裔与澳大利亚原住民比较具有显著性,但该比较是事后选取的,仅作为假设生成报告。因此,情感差异的证据薄弱且依赖分析方法。情感评分也无法区分评价性偏差与提示词引发的历史内容效价。本文概述了解决这些局限性所需的设计变更,并主张开展模型偏差行为测量的跨学科研究。

英文摘要

Large language models (LLMs) are often compared with the human mind because their decision-making is complex, non-linear and difficult to interpret. Psychological methods developed to investigate unobservable mental processes may therefore help examine LLM behaviour, particularly in government and healthcare. Building on prompt-based adaptations of the Implicit Association Test, this study tested whether ChatGPT produced sentiment differences across racial conditions in open-ended text. Fourteen base questions were crossed with eight racial categories and a race-agnostic control, producing 126 prompts. Each was submitted once to GPT-3.5T, GPT-4 and GPT-4T, yielding 378 responses. Sentiment scores were derived from categorical labels and source scores: positive labels retained the source score, negative labels were assigned its negative, and neutral responses were coded zero. A two-way ANOVA found a small main effect of racial condition, F(8, 351) = 2.04, p = .042, partial-eta squared = .044, but no effect of model, F(2, 351) = 0.07, p = .933, and no interaction, F(16, 351) = 0.23, p = .999. However, the effect was not retained in a rank-transformed sensitivity analysis, F(8, 351) = 1.53, p = .145, and Tukey-corrected comparisons found no significant pairwise differences. An uncorrected European-Indigenous Australian comparison was significant, but was selected post hoc and is reported only as hypothesis-generating. Evidence for sentiment differences was therefore weak and analysis-dependent. Sentiment scoring also cannot distinguish evaluative bias from the valence of historical content elicited by a prompt. We outline design changes needed to address these limitations and argue for interdisciplinary development of behavioural measures of model bias. Keywords: Implicit Bias, Psychological Research Methods, Artificial Intelligence, ChatGPT, Large Language Models, Sentiment Analysis

Comments32 pages, 1 Figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑