发表机构
Institute of Informatics and Telematics, National Research Council; Department of Computer Science, University of Pisa(国家研究委员会信息与远程信息处理研究所; 比萨大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究LLMs在伪装好与伪装坏条件下对暗黑三人格特质的表达调节,发现多数模型系统性改变得分,且情境和指令影响显著,凸显心理测量范式在评估反应扭曲中的价值。
AI 中文摘要
社会赞许性和印象管理是人类人格评估中反应扭曲的普遍来源,然而它们对大型语言模型(LLMs)的影响仍未得到充分探索。本研究调查当代LLMs是否在伪装好和伪装坏条件下系统性地调节暗黑三人格特质(马基雅维利主义、自恋和精神变态)的表达。研究在两种生态相关情境中评估了七个最先进的模型:就业选拔和法医评估,其中通过情境框架传达了社会赞许或非赞许的激励。特质表达使用标准心理测量评分程序进行测量,并在总体和项目层面与自我评估基线进行比较。结果揭示了系统性和条件一致的反应调节。大多数模型在伪装好条件下降低了暗黑三人格得分,在伪装坏条件下提高了得分,尽管这些效应的幅度和一致性在不同特质和模型间有所变化。马基雅维利主义和自恋表现出最强且最一致的转变,而精神变态则显示出更大的异质性。情境也影响了反应,就业场景通常比法医场景产生更大的效应。一项额外实验表明,明确的伪装坏指令比单纯的情境框架产生了显著更强的扭曲。结果表明,与人格相关的输出应结合其被引发的动机和情境背景来解读。更广泛地说,它们凸显了心理测量范式在评估对反应扭曲、印象管理和情境依赖行为转变的易感性方面的价值,对LLM基准测试、对齐评估和鲁棒性评估具有重要意义。
英文摘要
Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored. This study investigates whether contemporary LLMs systematically modulate the expression of Dark Triad traits (Machiavellianism, narcissism, and psychopathy) under fake-good and fake-bad conditions. Seven state-of-the-art models were evaluated across two ecologically relevant contexts: employment selection and forensic evaluation, in which socially desirable or undesirable incentives were conveyed through contextual framing. Trait expression was measured using standard psychometric scoring procedures and compared with self-assessment baselines at both aggregate and item levels. Results revealed systematic and condition-consistent response modulation. Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, although the magnitude and consistency of these effects varied across traits and models. Machiavellianism and narcissism showed the strongest and most coherent shifts, whereas psychopathy displayed greater heterogeneity. Context also influenced responses, with employment scenarios generally producing larger effects than forensic scenarios. An additional experiment showed that explicit fake-bad instructions generated substantially stronger distortions than contextual framing alone. The results suggest that personality-related outputs should be interpreted in light of the motivational and situational context in which they are elicited. More broadly, they highlight the value of psychometric paradigms for evaluating susceptibility to response distortion, impression management, and context-dependent behavioral shifts, with important implications for LLM benchmarking, alignment evaluation, and robustness assessment.
Comments21 pages, 7 figures, Journal