面向视觉语言模型空间推理的多视图关系蒸馏
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models
浏览论文内容
中文总结 AI 辅助
针对视觉语言模型空间推理的几何脆弱性问题,提出多视图关系蒸馏方法,在保留语言对齐的同时提升视觉空间推理性能,可泛化至3D场景理解任务。
中文摘要 AI 辅助
视觉语言模型(VLMs)在图像与视频理解领域已取得优异性能,但其视觉空间表征存在几何脆弱性问题,导致在具身智能、机器人技术及自动驾驶所需的空间推理任务中失效。现有几何对齐方法要么通过空间问答任务对VLMs进行微调,可能会固化虚假的视觉表征;要么融合大型几何对齐视觉模型的特征,这会大幅增加推理时的模型规模。从几何对齐视觉模型进行知识蒸馏是一种替代方案,但直接匹配多视图教师特征会破坏预训练的视觉与文本表征之间的对齐,进而损害对象及语言语义能力。本文提出多视图关系蒸馏(MVRD),该方法蒸馏跨视图的逐块余弦相似度而非教师特征本身。这些关系编码了足够用于空间理解的几何对应关系,同时使学生表征保持未完全确定的状态,从而使其能贴近自身预训练的视觉语言空间。在代表性VLMs上,MVRD提升了视觉空间推理能力,优于监督微调与特征蒸馏,且以显著更少的额外参数和更低的延迟接近特征融合方法。研究表明,MVRD在保留语言对齐的同时使视觉表征更具几何性,且可泛化到3D场景理解任务,如对象对齐、密集字幕及问答。
英文摘要
Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spatial reasoning needed for embodied AI, robotics, and autonomous driving. Prior approaches to geometry grounding either fine-tune VLMs on spatial question answering, which can perpetuate spurious visual representations, or fuse features from large geometry-grounded vision models, which substantially increases model size at inference. Knowledge distillation from geometry-grounded vision models offers an alternative, but directly matching multi-view teacher features can disrupt the pretrained alignment between visual and textual representations, degrading object- and language-semantic capabilities. We propose multi-view relational distillation (MVRD), which distills patch-wise cosine similarities across views instead of the teacher features themselves. These relations encode geometric correspondences adequate for spatial understanding, while leaving the student representation underdetermined, allowing it to remain close to its pretrained vision- language space. Across representative VLMs, MVRD improves visual-spatial reasoning, outperforming supervised fine-tuning and feature distillation while approaching feature fusion methods with considerably fewer added parameters and lower latency. We show that MVRD makes visual representations more geometric while retaining language alignment, and generalizes to 3D scene understanding tasks such as object grounding, dense captioning, and question answering.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。