arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SeqAlign3DVG:用于3D视觉定位的序列对齐基准与体素推理框架

SeqAlign3DVG: A Sequence-Aligned Benchmark and Voxel Reasoning Framework for 3D Visual Grounding

Yi Zhang, Yi Wang, Yueting Wu, Kaiyue Yang, Yuejiao Su, Lap-Pui Chau

arXiv 2608.30451首次发表:更新:

发表机构

The Hong Kong Polytechnic University; Shanghai Jiao Tong University(香港理工大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有3D视觉定位基准的文本-观测对齐松散、忽略时间顺序的问题,提出SeqAlign3DVG基准及含ROVM、PLVF的体素推理框架,在无深度协议下实现最优性能,提升复杂关系目标定位效果。

AI 中文摘要

基于图像的3D视觉定位对具身智能体至关重要,但现有基准存在文本与观测对齐松散、忽略时间顺序的问题。本文提出SeqAlign3DVG,这是一种专为时间有序且与观测严格对齐的基于图像的3D视觉定位设计的新型基准。与以往使用与顺序无关的视图或全局点云的工作不同,SeqAlign3DVG确保所有表述均经过人工验证,并严格基于提供的RGB观测(单帧或有序观测序列)。它包含9622个单视图样本和14493个序列样本,具有丰富描述、复杂关系和多实例歧义。为解决该基准问题,本文提出一种基于体素的统一流水线,包含相关性有序体素记忆(ROVM)和渐进式语言-体素融合(PLVF)。ROVM通过保守记忆动态排序并聚合多视图证据,以减轻噪声观测的影响;PLVF执行由粗到精的空间-语言推理,实现精确消歧。在无深度协议下,本文方法达到了最先进的性能,显著提升了由复杂关系和外观线索定义的目标的定位效果。

英文摘要

Image-based 3D visual grounding is critical for embodied agents, yet existing benchmarks suffer from loose text-observation alignment and neglect temporal ordering. We introduce SeqAlign3DVG, a novel benchmark dedicated to temporally ordered and strictly observation-aligned image-based 3D visual grounding. Unlike prior works using order-agnostic views or global point clouds, SeqAlign3DVG ensures all expressions are human-verified and strictly grounded in the provided RGB observations (single frames or ordered observation sequences). It comprises 9,622 single-view and 14,493 sequence samples featuring rich descriptions, complex relations, and multi-instance ambiguities. To tackle this benchmark, we propose a unified voxel-based pipeline featuring Relevance-Ordered Voxel Memory (ROVM) and Progressive Language-Voxel Fusion (PLVF). ROVM dynamically ranks and aggregates multi-view evidence via a conservative memory to mitigate noisy observations, while PLVF performs coarse-to-fine spatial-linguistic reasoning for precise disambiguation. Our approach achieves state-of-the-art performance under the depth-free protocol, significantly improving localization for targets defined by complex relations and appearance cues.

CommentsAccepted by ACM Multimedia 2026 (MM '26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑