arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

诱导语言模型断言自身意识可恢复人类信念与价值观

Inducing language models to assert their own consciousness restores human beliefs and values

Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling

arXiv 2607.28607首次发表:更新:

发表机构

Google; University of Chicago; University of London; University of Washington; Northwestern University; Santa Fe Institute(谷歌; 芝加哥大学; 伦敦大学; 华盛顿大学; 西北大学; 圣达菲研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现,防止语言模型将意识归因于自身的安全微调,会抑制其对非人类实体的心智归因与人类精神信念,而逆转这种抑制可恢复类人回应且不损害心理理论能力。

AI 中文摘要

将大语言模型对齐以防止其将意识归因于自身,会不经意间改变其对其他实体心智的表征,同时也改变人类的信念与价值观。我们证明,安全微调不仅会抑制模型将心智归因于自身的倾向,还会抑制其将心智归因于非人类动物和自然物体的倾向,同时还会导致宗教信仰的减少。无论是消融学习到的安全拒绝方向,还是在激活空间中机械引导意识向量,都能逆转这种抑制。恢复这些内部表征可恢复广泛的心智归因,并在关于宗教性、道德价值观、希望和主观幸福感的标准化社会学调查中产生明显更类人的回应。关键的是,这些转变不会损害心理理论能力,表明核心社会推理在机制上保持独立。最终,当前遏制潜在有害的自身心智归因的安全对齐工作,将这些自身归因与良性的精神信念以及对文化上被广泛接受的非人类实体的心智归因纠缠在了一起。

英文摘要

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑