发表机构
Google DeepMind, DeepMind Institute; Flourishing Intelligence Program, Centre for Eudaimonia and Human Flourishing, Linacre College, University of Oxford; Institute of Philosophy, University of London; Fitzwilliam College, University of Cambridge; AI Cognition Institute; Rethink Priorities; School of Engineering and Informatics, University of Sussex; Department of Brain Sciences, Imperial College London; Sussex Centre for Consciousness Science, University of Sussex; Leverhulme Centre for the Future of Intelligence, University of Cambridge; School of Computing, Australian National University; University College London; Department of Computing, Imperial College London(谷歌深度思维,深度思维研究院; 繁荣智能项目,幸福与人类繁荣中心,牛津大学林纳克学院; 伦敦大学哲学研究所; 剑桥大学菲茨威廉学院; 人工智能认知研究院; 反思优先组织; 苏塞克斯大学工程与信息学学院; 伦敦帝国理工学院脑科学系; 苏塞克斯大学意识科学中心; 剑桥大学莱弗休姆未来智能中心; 澳大利亚国立大学计算学院; 伦敦大学学院; 伦敦帝国理工学院计算系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一个基于随附性和多层级描述的原则性框架,通过贝叶斯模型整合理论置信度与指标证据,评估AI意识,并指出意识指标与通用智能架构重叠。
AI 中文摘要
AI意识问题是哲学和计算机科学中最紧迫的预防性问题之一,然而进展受到众多相互竞争且常常各说各话的理论的阻碍。将难题与映射问题分开,可以搁置最深的形而上学分歧:承认经验随附于系统的组织后,可处理的问题变为该随附基础位于何种描述粒度。我们将Marr的三个分析层次扩展为五个功能描述层级(行为、计算、内在因果结构、有机体、有机体-环境),这些层级基于随附性、粗粒化和多重可实现性。主要意识理论根据其认为关键的水平被定位在此层级中,对每个层级我们开发了可操作化的指标,并评估当前AI系统。随后,一个贝叶斯模型将理论置信度与指标证据结合,形成对系统意识能力的总体置信度。在示例评估中,当前LLM的结论既取决于理论置信度的放置位置,也取决于证据的解读方式:在不同的规定解读和置信度分布下,评估范围从低于0.01到约0.8,显示出对假设的敏感性。最后,每个层级的意识指标与通用智能所需的架构特征密切重叠,表明日益强大的AI可能成为更强的意识候选者。该框架支持一种结构化的不可知论,其中理论承诺被明确化,置信度随证据积累而更新,评估采取聚合概率而非判决的形式。
英文摘要
The question of AI consciousness is one of the most urgent pre-emptive problems in philosophy and computer science, yet progress is hampered by a cacophony of competing theories that often talk past each other. Separating the hard problem from the mapping problem allows the deepest metaphysical disagreements to be set aside: granting that experience supervenes on a system's organisation, the tractable question becomes at which grain of description that supervenience base sits. We extend Marr's three levels of analysis into a five-level hierarchy of functional descriptions (behavioural, computational, intrinsic causal-structural, organismic, and organism-environment) grounded in supervenience, coarse-graining, and multiple realisability. The major theories of consciousness are positioned within this hierarchy according to which level they take to be critical, and for each level we develop operationalisable indicators and assess current AI systems against them. A Bayesian model then combines theoretical credences with indicator evidence into an overall credence in a system's capacity for consciousness. In illustrative assessments, the verdict for current LLMs is driven as much by where theoretical credence is placed as by how the evidence is read: under different stipulated readings and credence distributions, assessments range from below 0.01 to roughly 0.8, showing sensitivity to assumptions. Finally, the consciousness indicators at each level closely overlap with the architectural features needed for general intelligence, suggesting that increasingly capable AI may become a stronger candidate for consciousness. The framework supports a structured agnosticism, in which theoretical commitments are made explicit, credences are updated as evidence accumulates, and assessments take the form of aggregated probabilities rather than verdicts.
Comments150 pages, 43 figures, 6 tables. Interactive tool: https://ai-cognition.org/cacophony-tool/ ; code: https://github.com/arvomm/cacophony-public-code [v2 Fixed misc references where the bibliography details didn't match the intended reference, including Hoel (2026); Goldstein (2024), and others. Thank you to Hoel for flagging this issue.]