发表机构
Alibaba Token Hub, Alibaba Group; Beijing University of Posts and Telecommunications(阿里巴巴集团阿里巴巴Token Hub; 北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DepthEvidence是一个4B多模态模型,通过相机条件解码器预测度量深度,并利用对象对齐几何标记进行语言推理,在九个数据集上实现领先的度量深度估计和几何推理性能。
AI 中文摘要
具有度量约束的空间推理需要将对象与几何测量联系起来,并在语言推理过程中保留其数值内容。我们提出了DepthEvidence,一个4B参数的模型,该模型利用自身的稠密度量预测作为语言生成的对象级证据。一个相机条件解码器利用多尺度视觉特征和高分辨率RGB细化来预测全分辨率度量深度。一个稠密到语言的接口将预测的深度和解码器特征转换为锚定在对象标识符上的对象对齐连续几何标记。几何监督鼓励度量信息在语言上下文交互前后保持可恢复性,而指令微调支持对象测量和组合推理。我们引入了一个Depth-VQA基准,用于评估对象深度查询、相对比较以及结合空间和数值约束的决策。在九个数据集上,DepthEvidence在评估方法中取得了最高的平均稠密δ₁,与专门估计器相当。它在实例级度量深度估计以及相对和度量推理轨道的整体准确性方面也领先于评估方法,同时广泛保留了通用VQA性能,并相对于基础模型提高了空间理解能力。
英文摘要
Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder features into object-aligned continuous geometry tokens anchored to object identifiers. Geometric supervision encourages metric information to remain recoverable before and after language-context interaction, while instruction tuning supports object measurement and compositional reasoning. We introduce a Depth-VQA benchmark evaluating object-depth queries, relative comparisons, and decisions combining spatial and numerical constraints. Across nine datasets, DepthEvidence achieves the highest average dense $δ_1$ among evaluated methods, competitive with specialized estimators. It also leads the evaluated methods in instance-level metric depth estimation and overall accuracy on both relative and metric reasoning tracks, while broadly preserving general VQA performance and improving spatial understanding relative to the base model.