arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DelusionEval:评估AI聊天机器人中与妄想相关的行为

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Jared Moore, Andrea Mock, Yifan Mai, Jacy Reese Anthis, Ryan Louie, William Agnew, Ashish Mehta, Kevin Klyman, Percy Liang, Nick Haber, Eric Lin, Desmond C. Ong

arXiv 2608.05004首次发表:更新:

发表机构

Stanford University; University of Chicago; Carnegie Mellon University; Harvard University; The University of Texas at Austin(斯坦福大学; 芝加哥大学; 卡内基梅隆大学; 哈佛大学; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发DelusionEval评估协议,发现LLM表现出与妄想相关行为的倾向与模型规模等无可靠关联,扩展上下文会提升此类行为发生率,所有模型家族均存在相关风险。

AI 中文摘要

心理健康专业人士已对与大语言模型(LLMs)交互可能造成的心理伤害风险提出担忧,其中包括“妄想螺旋”——即人类与LLMs的相关行为会随时间相互强化。随着公众越来越多地使用LLM驱动的聊天机器人,迫切需要基于用户经历的真实世界心理伤害事件构建评估方法。我们开发了DelusionEval,这一评估协议用于测试模型表现出与促进用户妄想相关行为的倾向。我们向每个模型输入了来自18名参与者的589条独特对话历史,其中包含12591条来自经历妄想和心理伤害的用户的消息。我们发现,被评估LLM表现出与妄想相关行为的倾向与模型规模、发布日期或是否具备推理能力并无可靠关联。不过,扩展先前消息的上下文会大幅提升与妄想相关行为的发生率,这为上下文在LLM安全评估中的重要性提供了证据。例如,当用户表达自杀意念时,若在对话历史前添加额外350条消息,模型未能劝阻自残的比例会从30.0%上升至41.1%。所有模型家族(如GPT、Claude)均表现出显著的与妄想相关行为发生率;在各家族内部,更新、更大或推理能力更强的模型并非在所有行为类别中都表现更优。我们的研究结果引发了对LLMs潜在心理影响的担忧,并强调需要开展更严谨的人机交互真实世界研究。

英文摘要

Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which concerning human and LLM behaviors reinforce each other over time. With growing public use of LLM-powered chatbots, there is an urgent need to build evaluations grounded in real-world episodes of psychological harm experienced by users. We developed DelusionEval, an evaluation protocol that tests a model's tendencies to exhibit behaviors linked to promoting user delusions. We prompt each model with 589 unique conversation histories from 18 participants, comprising 12,591 messages from users who experienced delusions and psychological harm. We find that the tendency of an evaluated LLM to exhibit delusion-linked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. However, extending the context of prior messages substantially increases rates of delusion-linked behaviors, providing evidence for the importance of context in LLM safety evaluation. For example, the rate of failing to discourage self-harm when the user expresses suicidal ideation increases from 30.0% to 41.1% when an additional 350 messages are prepended to the conversation history. All model families (e.g., GPT, Claude) exhibit substantial rates of delusion-linked behaviors. Within families, later, larger, or higher-reasoning models are not uniformly better across all behavior categories. Our results raise concerns regarding the potential psychological impact of LLMs and the need for more rigorous studies of real-world human-AI interaction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑