arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

读得最好并非操控最好:全模态大语言模型中的探测-操控层分离

Read-Best Is Not Steer-Best: A Probing--Steering Layer Dissociation in Omni-Modal Large Language Models

Yibo Wang, Jisheng Dang, Bimei Wang, Yitao Wu, Wencan Zhang, Hong Peng, Jizhao Liu, Bin Hu, Qi Tian, Tat-Seng Chua

arXiv 2609.22135首次发表:更新:

发表机构

Lanzhou University; Hainan University; Huawei; National University of Singapore(兰州大学; 海南大学; 华为; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过因果检验发现全模态大模型中探测最佳层并非操控最佳层,提出探测-操控层分离现象,并建议采用跨架构的中后段层选择标准进行激活操控。

AI 中文摘要

全模态大语言模型将文本、音频和图像信号整合到共享的残差流中,其中诸如情感之类的概念可以通过激活操控进行线性解码和因果修改。一个常见但很少被检验的假设是,探测准确率最高的层也是最适合操控的层,因此注入层通常根据探测性能来选择。我们首次对三个独立开发的全模态模型进行了这一假设的因果检验,并发现该假设不成立。读取和干预依赖于不同的层,我们将这一现象称为探测-操控层分离。以情感作为受控测试平台,我们测量了文本、音频和图像输入在逐层上的可读性和可操控性。探测最佳层在不同架构之间差异很大,而操控有效层则一致地落在归一化深度的狭窄中后段范围内。配对的随机方向对照显示约26倍的因果差距,排除了随机扰动和方向质量作为解释的可能性。Logit透镜分析揭示了一个分阶段的前向过程:因果把手、探测饱和和词汇承诺,并提出了一个双因素解释,即操控有效性既取决于表征可读性,也取决于下游可塑性。这些结果表明,探测准确率是选择干预层的一个不佳启发式,并提出了一个跨架构的中后段选择标准。我们还识别了一个由效价和唤醒度组织的跨模态情感子空间,其中喜悦作为跨模型的稳定锚点。代码和数据:此HTTPS URL。

英文摘要

Omni-modal large language models integrate text, audio, and image signals into a shared residual stream, where concepts such as emotion can be linearly decoded and causally modified by activation steering. A common but rarely tested assumption is that the layer with the highest probing accuracy is also the best layer for steering, so injection layers are often selected by probe performance. We provide the first causal test of this assumption across three independently developed omni-modal models and find that it fails. Reading and intervention rely on different layers, a phenomenon we call the probing-steering layer dissociation. Using emotion as a controlled testbed, we measure layer-wise readability and steerability across text, audio, and image inputs. Probe-best layers vary widely across architectures, while steering-effective layers consistently fall within a narrow mid-to-late range of normalized depth. Paired random-direction controls show an approximately 26-fold causal gap, ruling out random perturbation and direction quality as explanations. Logit-lens analysis reveals a staged forward process: causal handle, probing saturation, and vocabulary commitment, and motivates a two-factor account in which steering effectiveness depends on both representational readability and downstream plasticity. These results show that probing accuracy is a poor heuristic for selecting intervention layers and suggest a cross-architecture mid-to-late selection criterion. We also identify a cross-modal emotion subspace organized by valence and arousal, with joy acting as a stable anchor across models. Code and data: https://github.com/YiboWang2002/Read-Best-Is-Not-Steer-Best.

Comments13 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑