arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoordRefer:基于多视图图像的坐标感知三维视觉定位

CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images

Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen, Youyu Chen, Yanmin Wu, Jian Zhang

arXiv 2608.05569首次发表:更新:

AI 中文总结

本文提出CoordRefer框架,解耦坐标帧选择与坐标条件三维定位,在ScanRefer数据集上实现定位精度提升,性能优于坐标不可知基线及部分显式三维输入方法。

AI 中文摘要

基于多视图图像的三维视觉定位任务会先预测一个坐标帧以定义坐标系,再回归三维边界框完成定位。然而现有方法联合优化坐标帧选择与边界框回归,导致坐标相对的边界框模糊性,降低了定位性能。这种模糊性源于同一边界框在不同坐标帧下有不同数值表示,产生多个优化目标,最终生成无效的折中预测。为解决该挑战,本文提出CoordRefer,一种坐标感知框架,将坐标帧选择与坐标条件定位解耦。CoordRefer先选择参考帧定义坐标系,再基于该坐标系进行三维边界框预测。我们执行坐标感知的监督微调,以建立坐标帧选择和坐标条件边界框回归,随后采用基于三维IoU奖励的Group相对策略优化,使两个阶段与下游定位质量对齐。在ScanRefer数据集上,使用Qwen3-VL-2B模型时,CoordRefer较坐标不可知基线实现了Acc@0.25提升11%、Acc@0.5提升7%,其几何细化变体超越了使用显式三维输入的方法。

英文摘要

Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods jointly optimize coordinate frame selection and box regression, leading to coordinate-relative box ambiguity and degraded grounding performance. This ambiguity arises because the same box admits different numerical representations across coordinate frames, creating multiple optimization targets and yielding invalid compromise predictions. To tackle this challenge, we propose CoordRefer, a coordinate-aware framework that decouples coordinate frame selection from coordinate-conditioned grounding. CoordRefer first selects a reference frame to define the coordinate system and then conditions 3D box prediction on the coordinate system. We perform coordinate-aware supervised fine-tuning to establish coordinate frame selection and coordinate-conditioned box regression, followed by Group Relative Policy Optimization with 3D IoU-based rewards to align both stages with downstream grounding quality. On ScanRefer with Qwen3-VL-2B, CoordRefer achieves gains of 11% in Acc@0.25 and 7% in Acc@0.5 over the coordinate-agnostic baseline, while its geometrically refined variant surpasses methods using explicit 3D inputs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑