我们能否解读音频大语言模型的思维?一个可解读的多语言中间层工作空间
Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
浏览论文内容
中文总结 AI 辅助
本研究通过logit lens分析Qwen3-Omni模型,发现其中间层可在输出前清晰呈现音频问题答案,揭示了音频模型推理的特性与信号分布,为解读音频大语言模型内部思维提供了定性依据。
中文摘要 AI 辅助
音频大语言模型在特定方面是一个黑箱:我们只能看到它的输出,却无法知晓其推理过程,只有当模型将推理过程记录下来时,思维链监控才会发挥作用。我们使用logit lens(对数透镜)在音频token位置上分析基础模型Qwen3-Omni,发现模型在输出任何token之前,其中间层就已能以文字形式清晰呈现口语问题的答案。由此得出五项发现:(1)该读出内容不包含问题、选项或模型自身转录中的概念:在一段逐字转录为混乱噪音的音频片段中,它重构出“水门事件”和“丑闻”,经历“总统”角色,最终指向“尼克松”——这是一条隐藏的多步推理链,无需思维链即可读出;(2)内容与语言无关:一个由音频推断出的概念会同时以多种文字形式呈现,在英文输入中,38%的top-1(排名第一)读出内容为中文;(3)它是副语言的:将同一段音频作为输入,以及模型自身无情感的字幕,音频模型会形成字幕所缺失的声源、说话人角色或情感,且能更频繁地给出正确答案;(4)音频驱动的信号在输入时不存在,在网络约十分之一处开启,在中间层(深度的35%-80%)与文本先验最清晰地分离,激活修补实验表明,该信号在最后五分之一层之前就已被因果使用并确定;(5)删除单层可绘制出整个流程:声音读取集中在输入层,答案传递在输出层,而检索则分布在中间层。整个过程中,波形交换控制(文本相同,仅声音改变)将音频驱动信号与打印选项的先验隔离开来。这是对音频模型说话前所进行的推理的定性描述:所用指标为控制变量,而非基准分数。
英文摘要
An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni with a logit lens at the audio-token positions, we find that the answer to a spoken question becomes legible - in words - in the model's middle layers, before it emits any token. Five findings follow. (1) The readout carries concepts in neither the question, the options, nor the model's own transcription: on a clip whose verbatim transcription is empty garbling, it reconstructs Watergate and scandal, passes through the role president, and resolves to Nixon - a hidden multi-hop chain, read with no chain-of-thought. (2) The content is language-agnostic: one audio-inferred concept surfaces in several scripts at once, and 38% of top-1 readouts are Chinese on English inputs. (3) It is paralinguistic: given the same clip as audio and as the model's own emotion-free caption, the audio mind forms the sound source, speaker role, or affect that the caption discards, and answers correctly more often. (4) The audio-driven signal is absent at the input, turns on about a tenth of the way into the network, separates most cleanly from the text prior in the middle band (35-80% of depth), and activation patching shows it is causally used and committed before the last fifth of the layers. (5) Deleting single layers maps the pipeline: reading the sound in is localized to the entry layers and answer delivery to the output layer, while retrieval is distributed across the interior. Throughout, a waveform-swap control - identical text, only the sound changed - isolates the audio-driven signal from a prior over the printed options. This is a qualitative account of what an audio model works out before it speaks: the quantities are controls, not benchmark scores.
发表机构
- Amazon AGI Foundations(亚马逊AGI基金会)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。