发表机构
Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算学部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大型多模态模型推广到未见视觉模态的挑战,提出VVM-Tuning训练框架,通过模态合成和上下文让模型具备相关能力,引入VVM-Bench基准,实验证明经合成模态训练的模型在多模态上有改进且无需模态内训练。
AI 中文摘要
尽管大型多模态模型(LMMs)在RGB视觉方面取得了进展,但其推广到未见视觉模态的能力仍是一个未充分探索的挑战。我们认为不同视觉模态只是同一物理世界的不同采样。为此提出训练框架VVM-Tuning,通过模态合成和模态上下文使LMMs具备相关能力。具体包括从RGB场景合成多样图像训练模型解耦不变语义与变化外观,并将外观与语言对齐;在提示中引入模态上下文并通过指令调优辅助模型在推理时零样本适应未见模态。还引入VVM-Bench基准。实验表明经合成模态训练的5个测试模型在真实世界和新合成模态上均有一致改进且无需模态内训练。
英文摘要
Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue that different visual modalities are merely distinct samplings of the same physical world. Therefore, effective generalization requires models to possess both modality-agnostic perception of scene semantics and the adaptability to modality-specific characteristics. To achieve this, we propose a training framework, VVM-Tuning, to equip LMMs with these capabilities through modality synthesis and modality contexts. Specifically, we synthesize diverse appearance-varied images from RGB scenes, training the model to disentangle invariant semantics from varying visual appearances, and align these appearances with language for visual concepts decoupled from modalities. We then introduce modality contexts in the prompt and use instruction tuning to assist the model in mapping these appearance variations back to modality-related attributes, enabling zero-shot adaptation to unseen modalities during inference. To facilitate research in this direction, we introduce VVM-Bench, a comprehensive benchmark featuring 6 real and synthetic modalities to evaluate semantic perception and modality understanding. Experiments demonstrate that, via our training on synthetic modalities, 5 tested models exhibit consistent improvements on both real-world and novel synthetic modalities without in-modality training. Source code and data will be publicly available at https://github.com/Hunter-Will/VVM-Tuning.
CommentsAccepted by the European Conference on Computer Vision (ECCV) 2026