发表机构
Sony Group Corporation; Sony AI; University of North Carolina at Charlotte(索尼集团公司; 索尼人工智能; 北卡罗来纳大学夏洛特分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态大语言模型在特权模态设置下的问题,提出混合探针(MoP)框架及MoP跨模态训练(MoP-X),通过结构化探测机制分离模态信号,经评估在多任务中优于基线,有效利用辅助模态训练可提升性能。
AI 中文摘要
多模态大语言模型(MLLMs)通常假设训练期间可用的所有模态在推理时也可访问。但许多现实场景违反此假设,需要模型在特权模态设置下运行,即辅助模态仅在训练期间可用。现有MLLMs大多未能有效利用这些模态。我们提出混合探针(MoP)框架,通过结构化探测机制从共享模态编码器的中间表示中提取和组织信息,以分离特定模态和通用模态信号。还引入MoP跨模态训练(MoP-X),围绕探针解缠损失防止探针崩溃并鼓励跨模态学习。在针对特权模态设置的综合评估协议下评估MoP,其始终优于强大的MLLM基线,相对提升高达65%,证明有效利用辅助模态在训练时能带来显著收益。
英文摘要
Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, many real-world settings violate this assumption, requiring models to operate under a privileged modality setting, where auxiliary modalities are available only during training. While these modalities contain valuable information, existing MLLMs largely fail to leverage them effectively, as they treat modalities as interchangeable inputs rather than sources of complementary supervision. We propose Mixture of Probes (MoP), a novel framework that disentangles modality-specific and modality-general signals within the MLLM, allowing the model to preserve modality-dependent structure while learning transferable representations across modalities. At its core, MoP achieves this through a structured probing mechanism that extracts and organizes information from intermediate representations of a shared modality encoder, rather than relying only on final-layer alignment as done in existing MLLMs. To support this disentanglement, we further introduce MoP Cross-modal Training (MoP-X), a training strategy for MoP centered around a probe disentanglement loss that prevents probe collapse and encourages cross-modal learning. We evaluate MoP across two domains spanning eight tasks and four modalities under a comprehensive evaluation protocol tailored to the privileged modality setting, where each modality is independently treated as the sole input at inference time. MoP consistently outperforms strong MLLM baselines, achieving up to 65% relative improvement, demonstrating that auxiliary modalities, even when unavailable at inference, can provide substantial gains when effectively leveraged during training. Code, model checkpoints, and evaluation protocols will be made available at https://github.com/Sony/MoP.
CommentsPreprint (16 pages)