可读、忠实、可用:语言模型中人口统计身份的三个可分离属性
Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model
浏览论文内容
中文总结 AI 辅助
该研究探究LLM中人口统计身份的三个可分离属性,通过对Mistral-7B等模型的分析,发现可读、忠实、可用三者分离,为LLM模拟人群的相关辩论提供了新视角。
中文摘要 AI 辅助
大型语言模型被广泛用于模拟调查受访者,但其回答具有同质性,与真实的群体间差异不符。我们探究人口统计群体身份在大型语言模型(LLM)中的位置、其几何结构忠实反映真实群体间意见结构的程度,以及模型是否会利用其编码的身份信息。通过对Pew基准真值(ground truth)进行表征相似性分析,针对169个人口统计单元,我们对Mistral-7B模型中的1089个读出位置进行评分,并在6种属性类型上进行因果干预,得到四个结果:(1)标准的最后一个标记残差读出低估了模型的表现:在6种类型中的5种,注意力头读出的表现优于它,经选择校正后的保真度最高达rho=0.63——约为测量可靠性上限的70%——且在词汇相似性控制下仍保持有效;(2)单个注意力头(L11 H16)作为固定位置在全部6种类型中均显著忠实,而基于种族的类型则表现较弱且对提示敏感。这两种现象均得到复现:在第二个模型家族的三个检查点中,类似的注意力头在6种类型中的5种显著,且在同一种种族类型上表现最弱,即使经过百亿训练令牌,该模式几乎未发生变化;(3)因果使用与保真度并不一致:最清晰的因果通路位于保真度最低的类型之一(p=0.002,聚类稳健,固定深度),最保真的类型未呈现经校正的单层效应,替换全部身份信息仅使预测误差减少不到2%;(4)针对该单个注意力头的128维探测,其结果比模型自身答案更接近调查真值21%-31%——但几乎未恢复任何问题层面的群体排序,表现与模型自身答案无显著差异。可读、忠实排列、可因果使用是同一模型的三个可分离属性;将三者视为单一主张,是导致“大型语言模型能否模拟人群”这一辩论悬而未决的原因。
英文摘要
Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or uses has remained untested. Using representational similarity analysis against Pew American Trends Panel ground truth, we score demographic read-out locations in Mistral-7B and intervene causally across six attribute types. The internal geometry is faithful: attention-head read-outs dominate the standard residual read-out, reaching selection-corrected $ρ$ up to 0.63 -- about 70% of the measurement-reliability ceiling -- and one head, L11 H16, is significantly faithful across all six types, though race-based types stay weak and prompt-fragile, replicating in a second model family. Yet causal use does not track fidelity: the clearest causal pathway ($p=0.002$) sits in one of the least faithful types, the most faithful type shows no correction-surviving effect, and full identity swaps in the prompt move predictions by under 2% of their error. A 128-dimensional probe on that head lands 21-31% closer to survey truth than the model's answers, yet recovers almost none of the per-question group ordering. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the "can LLMs simulate populations" debate unresolved.