arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35036cs.AI

人格跟随并非选择性控制:大语言模型用户模拟中的中性差距

Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation

  • City University of Hong Kong(香港城市大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiashen Ren, Wenlin Zhang, Bohan Zhang, Xiaopeng Li, Zichuan Fu, Wanyu Wang, Junyi Li, Xiangyu Zhao

AI总结:

本研究揭示大语言模型用户模拟中人格提示的跨属性影响,提出中性差距概念,并设计三状态诊断法证明成功的人格跟随并不保证选择性控制。

AI中文摘要:

人格提示(Persona prompting)被广泛用于使用大语言模型(LLMs)构建用户模拟,然而它依赖于一个在很大程度上未经检验的假设:指定一个用户属性应仅改变该属性。我们检验了这一假设,并识别出选择性控制的一种系统性失败:在我们审计的所有八个黑盒大语言模型中,改变一个目标属性也会导致未指定的非目标属性上的响应发生偏移。例如,将用户描述为更倾向于冒险会改变颜色选择,尽管提示从未提及颜色;我们将此称为跨属性影响。语义、上下文和内部的分析共同表明,模型将人格提示视为关于用户的证据,并将推断出的画像扩展到未指定的偏好,我们将这一过程称为特质条件补全。接下来,我们询问明确指定非目标属性是否能恢复选择性控制。当非目标属性被赋予明确方向时,模型通常遵循声明并抑制目标属性的影响。然而,当同一属性被声明为中性时,目标属性继续影响所有五个开放权重检查点的选择,即使模型正确报告了声明的状态。这种差异,即中性差距,表明成功的人格跟随并不意味选择性人格控制,后者还要求保持非目标属性的稳定。我们用一种三状态诊断来操作化这一区别,该诊断将非目标属性保持未指定或声明为定向或中性;因为定向测试可以通过简单地遵循所述人格而通过,中性状态揭示了它们所遗漏的失败。在对独立条目的事后分析中,中性声明使51-81%的条目对目标敏感,而在定向声明下最多只有320个条目-极性比较中的1个。

英文摘要:

Persona prompting is widely used to construct user simulations with large language models (LLMs), yet it relies on a largely untested assumption: specifying one user attribute should change that attribute alone. We test this assumption and identify a systematic failure of selective control: across all eight black-box LLMs we audit, changing a target attribute also shifts responses on unspecified, non-target attributes. For example, describing a user as more risk-seeking shifts color choices, even though the prompt never mentions color; we term this cross-attribute influence. Semantic, contextual, and internal analyses collectively suggest that models treat a persona prompt as evidence about the user and extend the inferred profile to unspecified preferences, a process we call trait-conditioned completion. We next ask whether explicitly specifying non-target attributes restores selective control. When a non-target attribute is assigned a clear direction, models generally follow the declaration and suppress the target attribute's influence. However, when the same attribute is declared neutral, the target continues to affect choices across all five open-weight checkpoints, even when the model correctly reports the declared state. This disparity, the neutrality gap, demonstrates that successful persona following does not imply selective persona control, which additionally requires keeping non-target attributes stable. We operationalize this distinction with a three-state diagnostic that leaves the non-target attribute unspecified or declares it directional or neutral; because directional tests can be passed by simply following the stated persona, the neutral state reveals failures they miss. In a post hoc analysis of independent items, neutral declarations leave 51-81% of items target-sensitive, against at most 1 of 320 item-pole comparisons under directional ones.

补充信息

↑