发表机构
University of Bern; École Polytechnique Fédérale de Lausanne (EPFL); Idiap Research Institute; Machine Alignment, Transparency, and Security (MATS)(伯尔尼大学; 洛桑联邦理工学院; 伊迪亚普研究所; 机器对齐、透明性与安全研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI理解问题,本文提出复调概念:LLM输出源于并行机制的联合,并据此发展以健全电路为核心的理解观,以指导对AI的信任。
AI 中文摘要
当医生、法官或工程师必须决定是否信任AI模型的输出时,他们无法回避模型理解了什么这一问题。纯粹的数学或统计描述难以区分可信与不可信的输出,除非以另一种名义重新引入AI理解的问题。然而,这一问题目前的框架本身就不恰当,因为所继承的概念运作于单调范式之中:即一个认知系统对某事物的理解必须定位于支撑该理解所赋予的所有能力的单一机制。基于广泛的机制性证据,我们表明大语言模型(LLM)普遍具有复调性:输出产生于可靠性不均的并行机制的联合,这些机制以各种方式互补、重复或相互压制,且多个联合足以完成一项任务,而没有任何一个联合是不可或缺的。复调性不仅使理解的归因复杂化,还使单调推理模式变得危险。为此,我们发展了一种适合复调AI的理解概念。该概念以可靠的、被正确调用并控制输出的健全电路为核心。由此,理解的归因成为关于内部组织的可处理主张,并能够承担引导对AI信任的工作。
英文摘要
When a doctor, a judge, or an engineer must decide whether to trust an AI model's output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs without reintroducing the question of AI understanding in all but name. Yet the question is ill-framed as it stands, because the inherited concept operates within a monophonic paradigm: the idea that a cognitive system's understanding of something must be localised to a single mechanism underpinning all the capacities conferred by such understanding. Drawing on a wide range of mechanistic evidence, we show that LLMs are pervasively polyphonic: outputs emerge from coalitions of parallel mechanisms of uneven reliability, which variously complement, duplicate, or drown out one another, with several coalitions sufficing for a task without any one being indispensable. Polyphony not only complicates attributions of understanding, but renders monophonic inference patterns hazardous. In response, we develop a conception of understanding fit for polyphonic AI. It centres on sound circuitry that is reliably and correctly recruited and in control of outputs. Attributions of understanding thereby become tractable claims about internal organisation, and can do the work of guiding trust in AI.