发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对视觉语言模型空间推理短板,提出GPD方法,将几何证据作为特权融入在线策略自蒸馏,在4B主干模型上于多空间推理基准取得优于GRPO等方法的成绩。
AI 中文摘要
空间推理仍是视觉语言模型(VLMs)的长期短板,因为RGB输入无法直接提供几何证据。现有解决方案要么在推理时向模型注入3D信息,付出架构与延迟成本;要么用仅监督最终答案的结果奖励训练。空间错误源于感知:误判的深度或方向只能通过场景真实几何修正,而空间训练语料库的3D扫描源已具备该几何信息。本文提出GPD(Geometry-Privileged Distillation,几何特权蒸馏),将几何证据作为在线策略自蒸馏(OPSD)中的特权信息。针对每个问题,深度、语义和鸟瞰图(BEV)线索被渲染为紧凑文本,与参考答案一同路由至教师模型;仅对错误轨迹应用的特权KL散度增强了GRPO,部署的模型仍仅使用RGB输入。在4B主干模型上,GPD在VSI-Bench上取得57.1的成绩,在MindCube、SPARBench、MMSI-Bench和ViewSpatial上的平均分为37.6,在所有空间推理基准上均优于GRPO和答案特权型OPSD。 ablation实验证实了3D特权与答案特权的互补性、问题条件路由相比全上下文注入的优势,以及将蒸馏限制在错误轨迹的益处。
英文摘要
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying architecture and latency costs, or train with outcome rewards that supervise only the final answer. Spatial errors originate in perception: a misjudged depth or direction can be corrected only by the scene's true geometry, which the 3D-scanned sources of spatial training corpora already provide. We propose GPD (Geometry-Privileged Distillation), which makes geometric evidence the privilege in on-policy self-distillation (OPSD). For each question, depth, semantic, and bird's-eye-view (BEV) cues are rendered as compact text and routed to the teacher alongside the reference answer; a privileged KL, applied only to incorrect trajectories, augments GRPO, and the deployed model remains RGB-only. On the 4B backbone, GPD achieves 57.1 on VSI-Bench and 37.6 average across MindCube, SPARBench, MMSI-Bench, and ViewSpatial, outperforming both GRPO and answer-privileged OPSD across spatial reasoning benchmarks. Ablations confirm the complementarity of 3D and answer privilege, the advantage of question-conditioned routing over full-context injection, and the benefit of restricting distillation to incorrect trajectories.
CommentsCode available at https://github.com/ZJU-REAL/GPD