定位并控制大语言模型中的隐式个性化
Locating and Controlling Implicit Personalization in Large Language Models
查看机构详情
- Indiana University(印第安纳大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对5种大语言模型,发现追踪推荐变化的局部内部激活信号与隐式个性化行为相关,移除对应信号可抑制线索影响且保留基准性能,为控制大语言模型隐式个性化提供了因果控制思路。
中文摘要 AI 辅助
大语言模型(LLMs)即便在用户未明确说明人口统计身份时,也常因隐式人口统计线索而改变输出。此前研究已记录该行为,但这些行为变化与模型内部激活的关联仍不明确。我们针对5种LLMs开展匹配的带线索与中性对话研究,发现存在局部内部激活信号追踪推荐变化,相关系数最高达r=0.87。当多线索同时出现时,其内部信号大致结合,但输出变化并非简单相加。我们进一步表明,移除与某一线索相关的内部信号可抑制其影响,效果通常优于通过提示要求模型忽略人口统计信息,且能在很大程度上保留通用基准性能。不过,选择性移除某一维度影响同时保留共存维度的能力仍高度依赖模型与属性。这些结果将隐式个性化行为与可分析、可因果控制的内部信号关联起来。
英文摘要
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r=0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension's influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.