PGL-3D:面向三维视觉查询定位的渐进式几何学习
PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization
浏览论文内容
中文总结 AI 辅助
PGL-3D提出预测-选择-细化-重新预测框架,利用中间立方体几何引导特征更新,提升3D视觉查询定位性能,stAP达0.270,并在GSOT3D上提升mAO至25.78%。
中文摘要 AI 辅助
三维视觉查询定位(3DVQL)在RGB-点云序列中检索查询对象的最新连续出现,并为每个响应帧预测一个9自由度(9-DoF)的立方体。查询是独立于搜索序列捕获的,因此其标注姿态可能与对象在搜索帧中的外观不同。基准基线在特征建模后预测立方体,未将其几何信息用于后续的特征细化。我们研究了完整的中间立方体是否能在最终解码前改进查询和提议表示。我们提出了面向3DVQL的渐进式几何学习(PGL-3D),这是一种预测-选择-细化-重新预测框架,利用中间立方体引导搜索证据的聚合,并更新查询和提议表示。一个共享头首先为每个提议预测一个完整的立方体。然后,查询-管-记忆(QTM)通过结合提议关联、立方体质量、帧响应和目标缺失来选择参考观测,因为仅有关联置信度既不能确定目标存在性,也不能确定几何准确性。每个选定立方体的中心、大小和方向定义了基于查询条件的提议特征的软池化权重。池化后的记忆更新查询和提议表示,头从更新后的特征重新预测。一个仅用于训练的目标ST-D9O通过向参数回归添加边界、符号距离和软重叠项来监督每个阶段的立方体几何。PGL-3D在3DVQL上实现了平均stAP为0.270±0.004,而LaF报告的为0.044。消融实验支持了几何引导特征更新的益处,而分阶段分析显示立方体准确性有所提高。在我们复现的PROT3D中,用ST-D9O替换几何目标将GSOT3D上的mAO从21.63%提高到25.78%。我们的代码和模型将发布。
英文摘要
3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search frames. The benchmark baseline predicts cuboids after feature modeling, leaving their geometry unused for subsequent feature refinement. We investigate whether complete intermediate cuboids can improve query and proposal representations before final decoding. We introduce Progressive Geometric Learning for 3DVQL (PGL-3D), a predict--select--refine--re-predict framework that uses intermediate cuboids to guide the aggregation of search evidence and update query and proposal representations. A shared head first predicts a complete cuboid for every proposal. Query--Tube--Memory (QTM) then selects reference observations by combining proposal association, cuboid quality, frame response, and target absence, since association confidence alone establishes neither target presence nor geometric accuracy. The center, size, and orientation of each selected cuboid define soft pooling weights over query-conditioned proposal features. The pooled memory updates the query and proposal representations, and the head re-predicts from the updated features. A training-only objective, ST-D9O, supervises cuboid geometry at every stage by adding boundary, signed-distance, and soft-overlap terms to parameter regression. PGL-3D achieves a mean stAP of $0.270 \pm 0.004$ on 3DVQL, compared with $0.044$ reported for LaF. Ablations support the benefits of geometry-guided feature updates, while stage-wise analyses show improved cuboid accuracy. Replacing the geometry objective in our PROT3D reproduction with ST-D9O improves mAO on GSOT3D from $21.63\%$ to $25.78\%$. Our code and models will be released.
发表机构
- Wuhan University(武汉大学)
- University of North Texas(北德克萨斯大学)
- Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
机构由 AI 辅助整理,请以论文原文为准。