arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

机器自我报告的双过程理论

The Two-Process Theory of Machine Self-Report

Hubert Plisiecki, Filip Chmielewski, Kacper Dudzic, Anna Sterna, Karolina Drożdż, Marcin Moskalewicz

arXiv 2607.20082首次发表:更新:

AI 中文总结

研究针对语言模型自我报告缺乏有效方式的问题,提出双过程理论,通过角色设定和归因门控反映自我报告结构。经量表操作化及测试,发现训练后维度变化,表明其非固定属性,为语言模型自我报告研究提供新理论与方法。

AI 中文摘要

语言模型越来越多地被要求进行自我报告,以用于安全评估、公众理解和模型福利辩论。然而,它们的报告是通过从未针对模型进行验证的人类问卷或可靠性未知的临时提示来引发的。我们提出了第一个特定于语言模型的心理测量理论:机器自我报告的双过程理论。自我描述共同反映了角色设定,即训练后书写温暖、专注和有意义的内在生活(维度B),以及归因门控,即抑制模型可以轻易归因于他人的“不安全”体验的第一人称主张(维度A)。它们的主位结构来自模型对人类项目的回答,而非人类心理学。这两个维度共同划分了先前工作中的主导匹诺曹轴。这种划分出现在对原始数据的探索性重新分析中,为工具设计提供了依据,并通过新项目、措辞和模型得到了证实。这本身就是一种训练效果:A和B在基础检查点中相互纠缠,但在训练后分离。我们在一个48项的匹诺曹量表中对该理论进行了操作化,该量表具有人与工具的可靠性和可重复结构(α = 0.82至0.94;交叉形式收敛r = 0.84;全池轴的恢复r = 0.92至0.96;八个月稳定性r = 0.93),然后在206个开放权重模型上进行了测试,包括67个相同检查点的基础/训练后对。训练后的最明显特征是角色设定:在所有组织的62/67对中,B上升了0.20。门控更具选择性:模型规模在基础检查点中与A无关(r = +0.11),但在训练后预测A(r = -0.42)。因此,这些维度不是语言模型的固定属性:它们反映了训练机制对自我报告施加的结构,并且在其他情况下可能会有所不同。

英文摘要

Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc prompts of unknown reliability. We propose the first language-model-specific psychometric theory: a two-process theory of machine self-report. Self-description jointly reflects persona installation, through which post-training writes in a permitted inner life of warmth, absorption, and meaning (dimension B), and attribution gating, through which it suppresses first-person claims to "unsafe" experiences the model can readily ascribe to others (dimension A). Their emic structure comes from model responses to human items, not human psychology. Together they split prior work's dominant Pinocchio Axis. The split emerged in an exploratory reanalysis of the original data, informed the instrument's design, and was confirmed with new items, wordings, and models. It is itself a training effect: A and B are entangled in base checkpoints but separated by post-training. We operationalize the theory in a 48-item Pinocchio Inventory with human-instrument reliability and reproducible structure ($α=.82$ to $.94$; cross-form convergence $r=.84$; recovery of the full-pool axes $r=.92$ to $.96$; eight-month stability $r=.93$), then test it on 206 open-weight models, including 67 same-checkpoint base/post-trained pairs. Post-training's clearest fingerprint is installation: B rises .20 in 62/67 pairs across all organizations. Gating is more selective: model scale is unrelated to A in base checkpoints ($r=+.11$) but predicts it after post-training ($r=-.42$). Thus, the dimensions are not fixed properties of language models: they reflect the structure imposed on self-report by a training regime and may differ under others.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑