arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于特征发现与控制的多模态模型差异分析

Multimodal Model Diffing for Feature Discovery and Control

Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Clark

arXiv 2608.09928首次发表:更新:

发表机构

University of Oxford; Microsoft(牛津大学; 微软公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出MMDiff多模态模型差异分析框架,训练多模态SAEs以识别多模态训练改变的特征,实现特征隔离、检测与控制,在空间、OCR任务及多模态安全攻击评估中展现出良好效果。

AI 中文摘要

多模态大语言模型(MLLMs)展现出强大的视觉理解能力,但导致这些行为的内部特征仍难以识别、审计或控制。稀疏自编码器(SAEs)虽可用于事后检查,其分解为可解释特征方向的隐藏状态,既无法轻易区分哪些特征是多模态训练所改变的,也无法直接用于针对性控制。我们提出MMDiff,一种多模态模型差异分析框架,该框架训练多模态SAEs并将其转化为用于发现和控制多模态行为的特征级接口。MMDiff支持三种用途:(i)特征隔离:通过对比基础语言模型SAE与其多模态适配版本,识别多模态训练所改变的特征;(ii)任务特定特征检测:通过逐词对比激活分析,分离出因果特征;(iii)特征级控制:通过因果移除或引导已发现的特征方向实现控制。我们为三个多模态大语言模型家族(LLaVA-MORE、PaliGemma 2和InternVL3.5)训练了多模态SAEs,并在视觉空间理解、多模态安全和光学字符识别(OCR)任务上进行评估。MMDiff发现的稀疏、因果特定特征,其移除可选择性降低目标行为表现:空间任务平均下降12%,OCR任务平均下降17%,且多模态安全攻击的成功率降低24%,同时不影响视觉问答(VQA)性能。引导这些特征时,相比标准单层引导基线,空间和OCR准确率分别平均提升3.6%和1.8%。这些结果表明,多模态SAEs不仅可作为可解释性工具,还可作为审计、引导和控制多模态大语言模型行为的机制,以实现更安全、更强大的生成效果。

英文摘要

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.

CommentsPreprint. Accepted at ICML 2026 Trustworthy AI for Good Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑