蒸馏RGB可恢复的内容:面向仅RGB视觉-语言模型的特权3D证据
Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出特权3D证据蒸馏框架,将训练时的3D证据转化为仅RGB视觉-语言模型的空间推理能力,使仅RGB学生模型在11项指标上均优于基线,验证了该方法的有效性。
中文摘要 AI 辅助
3D场景理解需要推理实体存在性、空间布局和物体关系,但仅RGB图像通常提供的3D线索不足。现有3D视觉-语言模型(3D-VLMs)在推理时通常依赖深度或感知3D位置的输入,这会带来额外的采集、重建或标注成本,限制了仅RGB部署的可能性。因此,本研究探讨如何将训练时的3D证据转化为仅RGB推理时仍能保留的空间推理能力。我们提出一种特权证据蒸馏框架,该框架通过统一证据接口和受控残差注入构建可蒸馏的教师模型,并通过logit和结构化表示蒸馏将其知识迁移给仅接收RGB图像和问题的可部署学生模型。为避免模仿RGB无法支持的教师信号,我们进一步引入证据敏感性引导的蒸馏,该方法使用损坏的证据识别高度依赖证据的目标并降低其监督权重。我们还基于匹配的基线、教师和学生定义了可恢复性分解,将特权增益分为RGB可恢复的改进和残差教师优势。在四个基准测试中,教师模型在对比方法的11个报告指标中7个取得最佳结果;仅RGB学生模型在全部11个指标上均优于其匹配的基线,包括ScanQA CIDEr提升10.4、Scan2Cap CIDEr@0.5提升19.1,且无需额外推理时输入。这些结果验证了训练时特权3D证据蒸馏对教师性能和可部署仅RGB空间推理的有效性。此外,我们的匹配基线-教师-学生分析表征了跨证据类型和空间技能的特权增益迁移。
英文摘要
3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient 3D cues. Existing 3D-VLMs commonly rely on depth or 3D-position-aware inputs at inference time, introducing additional acquisition, reconstruction, or annotation costs that limit RGB-only deployment. We therefore study how training-time 3D evidence can be converted into spatial reasoning capabilities retained under RGB-only inference. We propose a privileged-evidence distillation framework that constructs a distillable teacher through a unified evidence interface and controlled residual injection, and transfers its knowledge to a deployable student receiving only RGB images and questions through logit and structured representation distillation. To avoid imitating teacher signals unsupported by RGB, we further introduce evidence-sensitivity-guided distillation, which uses corrupted evidence to identify highly evidence-dependent targets and down-weight their supervision. We also define a recoverability decomposition based on the matched baseline, teacher, and student, separating privileged gains into RGB-recoverable improvements and residual teacher advantages. Across four benchmarks, the teacher achieves the best result on 7 of 11 reported metrics among the compared methods. The RGB-only student outperforms its matched baseline on all 11 metrics, including gains of 10.4 ScanQA CIDEr and 19.1 Scan2Cap CIDEr@0.5, without additional inference-time inputs. These results validate the effectiveness of training-time privileged 3D evidence distillation for both teacher performance and deployable RGB-only spatial reasoning. Separately, our matched baseline-teacher-student analysis characterizes privileged-gain transfer across evidence types and spatial skills.