Comments5 pages, 3 figures. Accepted at the 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence, and Networking (QPAIN)
机构
*
Fundation Model Research Center, CASIA(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
College of Automotive and Energy Engineering (CAEE), Tongji University(同济大学汽车与能源工程学院)
Inducing language models to assert their own consciousness restores human beliefs and values
诱导语言模型断言自身意识可恢复人类信念与价值观
Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling
机构
*
Google(谷歌)
;
University of Chicago(芝加哥大学)
;
University of London(伦敦大学)
;
University of Washington(华盛顿大学)
;
Northwestern University(西北大学)
;
Santa Fe Institute(圣达菲研究所)
CommentsThis manuscript supersedes the preliminary version available as arXiv:2511.18790. The work has been substantially revised, expanded, and reorganized, with a refined threat model, revised methodology, clearer stage-level evaluation criteria, and expanded analysis of moderation bypass, instruction reconstruction, and execution