arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DVD:用于高效基于多模态大语言模型(MLLM)感知的动态向量解码

DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception

Jinghua Hou, Zhe Liu, Hengshuang Zhao

arXiv 2610.12266首次发表:更新:

发表机构

The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对现有基于MLLM的感知方法的令牌开销大、范围精度受限问题,提出DVD动态向量解码方法,统一2D与3D感知表示,在多基准上实现更优性能并降低令牌开销与推理延迟。

AI 中文摘要

多模态大语言模型(MLLM)在连接视觉与语言方面已取得显著进展,为人类-机器交互、机器人技术及自动驾驶等领域所需的各类感知任务提供了便利。然而,现有基于MLLM的感知方法主要依赖基于文本的坐标表示,该方法存在令牌(token)开销过大的问题;或采用固定范围量化,该方法存在范围与精度限制,尤其在空间范围无界且对定位精度要求高的3D领域中问题更为突出。为应对这些挑战,本文提出一种名为DVD的动态向量解码方法,该方法统一了2D与3D感知任务的表示。具体而言,我们首先将各类感知表示(即2D边界框、2D掩码及3D边界框)转换为1D向量序列,随后将其映射至高维空间中的紧凑离散令牌。接着,一个轻量级反令牌器(de-tokenizer)通过将输出令牌解码回原始2D与3D感知表示,实现与MLLM的无缝集成。在RefCOCO系列、SUN-RGBD、KITTI、Hypersim、nuScenes等2D与3D感知基准上开展的大量实验表明,DVD在2D与3D任务中实现了更优性能,且大幅降低了令牌开销与推理延迟。DVD为将感知能力集成至MLLM提供了一种高效通用的框架,克服了现有方法的固有局限。

英文摘要

Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and reduces significantly the token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑