Qwen-3D:用于空间理解的通用三维视觉语言模型
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
浏览论文内容
中文总结 AI 辅助
本研究提出 Qwen-3D 三维视觉语言模型,通过多视角几何线索与三维旋转位置嵌入提升空间推理,在三维任务上超越现有 3D LMM,还保持二维任务性能。
中文摘要 AI 辅助
大型多模态模型(LMM)在图像和短视频上已取得显著成功,但由于基于帧的分词和有限的上下文窗口,将其扩展到长视频仍具挑战性。三维几何为视觉流提供了一种自然的压缩机制:深度和相机位姿可将多视角及时步的观测融合为持久的、与世界对齐的表示。尽管近期的三维大型多模态模型(3D LMM)利用几何感知表示提升了空间推理能力,但在定位和分割任务上仍落后于专业三维感知系统。我们认为一个关键局限是几何感知解码:现有方法通过语言 token、候选框选择或轻量定位查询传递三维预测,在语言推理与密集几何预测间形成瓶颈。基于这些见解,我们提出 Qwen-3D,一种几何感知的 LMM,它利用多视角几何线索在 Qwen 主干内压缩视觉信息,支持对静态场景的高效长程视觉推理。Qwen-3D 用三维旋转位置嵌入增强视觉 token,使注意力能直接在三维场景空间中运行,而非在独立图像帧间运行,从而促进可扩展的跨视角及时序推理。为弥合语言与几何的差距,Qwen-3D 采用基于查询的分割解码器,将语言直接定位到底层三维场景表示中,统一了图像和视频上的指代定位、实例分割及视觉问答。在多样的基准测试中,Qwen-3D 超越了现有三维大型多模态模型,且优于若干大型专有二维模型。值得注意的是,Qwen-3D 通过联合在二维和三维数据上训练,在取得这些改进的同时,保持了在标准二维视觉语言基准上的强性能。
英文摘要
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geometry-aware representations to improve spatial reasoning, they continue to lag behind specialist 3D perception systems on grounding and segmentation tasks. We argue that a key limitation is geometry-aware decoding: existing methods communicate 3D predictions through language tokens, proposal selection, or lightweight grounding queries, creating a bottleneck between language reasoning and dense geometric prediction. Building on these insights, we introduce Qwen-3D, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes. Qwen-3D augments visual tokens with 3D Rotary Positional Embeddings, allowing attention to operate directly in 3D scene space rather than across independent image frames and thereby facilitating scalable cross-view and temporal reasoning. To bridge language and geometry, Qwen-3D incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation, unifying referential grounding, instance segmentation, and visual question answering across both images and videos. Across a diverse set of benchmarks, Qwen-3D surpasses existing 3D LMMs and outperforms several large proprietary 2D models. Notably, Qwen-3D achieves these improvements while maintaining strong performance on standard 2D vision-language benchmarks by jointly training on 2D and 3D data.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。