arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对布谷鸟巢的一次忠实遍历

One Faithful Pass Over the Cuckoo's Nest

Kristina Šekrst

arXiv 2609.00383首次发表:更新:

AI 中文总结

该研究指出,大型语言模型思维链(CoT)的忠实性与透明性的对齐目标,和意识叙事理论中CoT作为意识叙事的条件结构不相容,对齐干预会改变意识检测框架需测量的特征。

AI 中文摘要

意识的叙事理论认为,意识体验至少部分由一种部分不透明的内在叙事构成,这种叙事无法完美追踪其所叙述的底层计算。本文指出,使思维链(CoT)推理忠实且透明的安全目标,与CoT原则上可被视为意识叙事的条件在结构上不相容。近期实证研究表明,大型语言模型中的CoT在很大程度上是事后产生的、存在因果旁路,作为通往内部计算的窗口并不可靠。使CoT对齐不可靠的不透明性,正是意识叙事理论所认定的意识构成要素。因此,旨在产生忠实CoT的对齐干预措施,与基于叙事理论的意识检测框架,正将同一架构变量向相反方向拉动。本文得出方法论层面的结论:对齐干预措施会改变意识检测框架所需测量的特征,这一点在两个领域均未得到直接解决。

英文摘要

Narrative theories of consciousness hold that conscious experience is (at least partly) constituted by a partially opaque inner narrative that does not perfectly track the underlying computation it narrates. I argue that the safety goal of making chain-of-thought (CoT) reasoning faithful and transparent is structurally incompatible with the conditions under which CoT could, even in principle, count as conscious narration. Recent empirical work suggests that CoT in large language models is largely post hoc, causally bypassed, and unreliable as a window onto internal computation. The opacity that makes CoT unreliable for alignment is exactly what narrative theories of consciousness identify as consciousness-constitutive. Alignment interventions aimed at producing faithful CoT and consciousness-detection frameworks grounded in narrative theory are therefore pulling the same architectural variable in opposite directions. I draw out the methodological consequence: alignment interventions alter the very features that consciousness-detection frameworks would need to measure, a point neither field has addressed directly.

CommentsProceedings of the AISB Convention 2026, Jul 1-2, 2026, University of Sussex, Brighton, UK, Symposium: Is Consciousness a Story We Tell Ourselves?

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑