arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更多模态总是有帮助吗?缺失模态鲁棒性的几何视角

Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

Songyuan Sui, Zhen Tan, Mohan Zhang, Rana Muhammad Shahroz Khan, Xia Hu, Tianlong Chen

arXiv 2610.04792首次发表:更新:

发表机构

Rice University; Stevens Institute of Technology; University of North Carolina at Chapel Hill(莱斯大学; 史蒂文斯理工学院; 北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文从几何视角揭示多模态模型在缺失模态时性能下降的机理,提出基于格拉斯曼子空间几何的轻量级参数编辑方法GU,通过沿测地线旋转子空间提升缺失模态鲁棒性,同时保持完整模态精度。

AI 中文摘要

缺失模态一直是多模态学习中的一个长期挑战。现有方法通常通过模态恢复或自适应策略来解决这一问题。然而,它们忽略了模型在多模态训练期间形成的内部跨模态依赖,这些依赖后来会损害鲁棒性。我们系统地刻画了一种违反直觉的部署时失效模式:当推理时缺少一种模态时,在完整模态上训练的模型可能表现不如单模态模型。这种模式出现在多种架构中,如融合模型、CLIP风格的双塔模型和视觉-语言模型。我们表明,这种退化与主参数子空间中学习到的跨模态依赖密切相关。多模态训练诱导了这些子空间的结构化旋转,特别是在跨模态交互层中。这些旋转与在缺失模态输入下任务对齐边际的减少和更大的任务感知表示损害相关。我们提出了测地线遗忘(Geodesic Unlearning, GU),一种轻量级参数编辑方法,利用格拉斯曼子空间几何进行结构化子空间校正,以提高缺失模态鲁棒性。它沿着测地线路径将主输入子空间向单模态参考旋转。我们证明,这种校正最小化了在固定子空间距离预算内到参考的距离。跨架构和数据集的实验表明,GU在缺失模态推理下提高了性能,同时保持了完整模态的准确性,优于强缺失模态鲁棒性基线。这些发现支持了部署时缺失模态退化的几何观点,并建议局部子空间编辑作为鲁棒性校正的实用途径。

英文摘要

Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It rotates the principal input subspace toward a unimodal reference along a geodesic path. We prove that this correction minimizes the distance to the reference within a fixed subspace-distance budget. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction.

CommentsNeurIPS 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑