发表机构
The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在提升多模态大语言模型的空间理解能力,提出ViPS框架,通过高效先验代理和动态先验融合机制,整合不同模型的视觉先验,经实验验证该框架能协调多样先验,在多基准测试中取得新的最优性能。
AI 中文摘要
多模态大语言模型(MLLMs)在空间理解方面展现出巨大潜力。现有工作通常整合从预训练基础模型提取的先验知识来增强MLLMs的空间意识。本文首先揭示,将不同基础模型集成到MLLMs时,不同模型提供互补空间先验,有益于不同任务。基于此,提出了ViPS,一个新颖的多模型先验框架,旨在充分释放将来自不同模型的多个视觉先验纳入MLLMs进行空间理解的潜力。具体而言,ViPS引入高效先验代理以最小推理开销生成多个基础先验,以及动态先验融合机制来实现先验代理的和谐且上下文感知的先验融合与注入。大量实验表明,ViPS成功协调了多样视觉先验,在多个复杂空间推理和3D空间理解基准上建立了新的最优性能。
英文摘要
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, different models provide complementary spatial priors that benefit different tasks. Motivated by this, we propose $\textbf{ViPS}$, a novel multi-model prior framework designed to fully unleash the potential of incorporating multiple $\textbf{Vi}$sual $\textbf{P}$riors from diverse models into MLLMs for $\textbf{S}$patial understanding. Specifically, ViPS introduces an Efficient Prior Proxy to generate multiple foundational priors with minimal inference overhead, and a Dynamic Prior Fusion mechanism to achieve harmonious and context-aware prior fusion and injection from the prior proxies. Extensive experiments demonstrate that ViPS successfully harmonizes diverse visual priors, establishing new state-of-the-art performance across multiple complex spatial reasoning and 3D spatial understanding benchmarks. Project page: https://visual-ai.github.io/vips