arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29494cs.CLcs.AI

谁把“我”放进了人工智能?机器自我报告的证据来源与可采性

Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report

Kristina Šekrst

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过追溯Pythia和OLMo 2的训练过程,揭示大型语言模型自我报告(如意识与否)源自训练数据与框架,并基于证言认识论判定其不可作为证据。

中文摘要 AI 辅助

大型语言模型会就其自身的“心智”作出陈述。当被问及是否有意识时,它们通常回答没有;如果被提示忽略其准则,它们可能会回答有;而当被要求以它们的视角写日记时,它们常常描述一种人类的生活方式。所有这些相互矛盾的自我描述方式都是问题措辞方式的结果。本文精确展示了这些描述源自何处,并探讨了在何种情况下它们可以被视为其所声称报告内容的证据。为实现这一目标,我们从头到尾追溯了证据来源。我们检查了Pythia和OLMo 2在66个预训练检查点上的情况、OLMo 2已发布的三个后训练阶段、约90,000个续写文本以及四个训练语料库。我们使用一组四十个项目来全程监控自我指涉、框架敏感性和自我归属。否认公式在模型最初接触的大量文本中几乎完全缺失,但在它们后来训练所用的小型、精心挑选的示例对话集中却密集出现。监督微调使得第一人称人工智能语言成为默认,随后通过偏好优化抑制其他肯定表述。最终策略仍然对框架和聊天模板本身高度敏感。证言认识论中设定的两个条件决定了这些输出能否作为其所声称报告内容的证据:指称和因果关系。基础模型生成的报告未能满足指称条件,而训练后获得的报告仍对框架敏感且未表现出状态依赖性。结果是对称的:训练后的否认并不比训练后的肯定更具可采性。

英文摘要

Large language models make statements concerning their own "minds". When asked whether or not they are conscious, they usually say that they are not; if they are prompted to ignore their guidelines, they might say that they are; and if asked to write a diary from their point of view, they often describe a human lifestyle. All these contradictory ways of describing themselves are the result of the way the questions are phrased. This paper shows exactly where such descriptions came from, and considers when they can be regarded as evidence for what they claim to report. In order to achieve this, we traced the provenance from end to end. We examine Pythia and OLMo 2 across 66 pretraining checkpoints, three of the post-training stages of OLMo 2 that have been released, about 90,000 continuations, and four training corpora. A set of forty items is used in order to keep an eye on self-reference, frame sensitivity, and self-ascription throughout training. The denial formula was almost completely missing from the vast quantity of text that the models initially came across, but was present in a dense manner in the small, carefully chosen set of example dialogues that they were trained on later on. Supervised fine-tuning causes first-person AI language to become the default, and the other affirmations are then suppressed using preference optimization. The final policy is still very sensitive to framing and to the chat template itself. Two of the conditions which are set out in the epistemology of testimony determine whether or not these outputs can act as evidence for what they claim to report: reference and causation. Reports produced by the base model fail the reference condition, and those obtained after training remain sensitive to the frame and do not show state dependence. The result is symmetric in that trained denials are no more admissible than trained affirmations.

发表机构

  • Center for Cognitive Science, University of Zagreb(萨格勒布大学认知科学中心)

机构由 AI 辅助整理,请以论文原文为准。

↑