AI 中文总结
针对Omni-LLMs存在的感知-决策失配问题,本文提出因果模态敏感性及对应诊断方法,构建CausalMSBench数据集,提出无需训练的MSA框架以恢复模型的因果模态敏感性。
AI 中文摘要
全模态大语言模型(Omni-LLMs)为世界动作模型、自主智能体等应用提供复杂多模态推理能力,但其出色性能常掩盖严重的感知-决策失配(PDM)问题,即决策与多模态感知不一致。为诊断该问题,本文提出因果模态敏感性(CMS),通过双视角框架实现:宏观行为层面的答案保留率(ARR),以及追踪微观分布变化的对数几率角差异(LAD);同时构建排除语言先验的诊断数据集CausalMSBench。基准测试显示,主流Omni-LLMs的CMS极低,移除关键模态时几乎无分布变化。为解决此问题,本文提出模态子空间激活(MSA),这是一种无需训练的推理时框架,通过奇异值分解(SVD)估计模态激活强度,动态平衡最后隐状态中的模态投影,有效恢复基准测试中的CMS。
英文摘要
Omni-Large Language Models (Omni-LLMs) power complex multi-modal reasoning in applications like World Action Models and autonomous agents. However, their strong performance often masks a profound Perceptual-Decision Misalignment (PDM), where decisions remain unfaithful to multi-modal perceptions. To diagnose this, we formalize Causal Modality Sensitivity (CMS), operationalized via a dual-lens framework: Answer Retention Rate (ARR) at the macro behavioral level, and Logit Angular Discrepancy (LAD) to track microscopic distribution shifts. We also curate CausalMSBench, a diagnostic dataset isolating language priors. Benchmarking reveals that popular Omni-LLMs exhibit critically low CMS, showing negligible distribution shifts even when key modalities are removed. To rectify this, we propose Modality Subspace Activation (MSA), a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths. MSA dynamically balances modal projections in the last hidden state, effectively restoring CMS across benchmarks.