arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Talk2Sensors:基于传感器自适应物理线索匹配的自动驾驶3D视觉 grounding

Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching

Runwei Guan, Di Tian, Ningwei Ouyang, Ruixiao Zhang, Shaofeng Liang, Haocheng Zhao, Lianqing Zheng, Xiaokai Bai, Guotao Wang, Daizong Liu, Henghui Ding, Hui Xiong

arXiv 2608.04568首次发表:更新:

发表机构

Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou); MMLab, CUHK; School of Transportation Science and Engineering, Harbin Institute of Technology; School of Advanced Technology, Xi’an Jiaotong-Liverpool University; School of Electronics and Computer Science, University of Southampton; School of Intelligent Manufacturing and Smart Transportation, Suzhou City University; Qingdao University of Science and Technology; Institute for Math & AI, Wuhan University; Institute of Big Data, Fudan University(香港科技大学(广州)人工智能学域; 香港中文大学MMLab; 哈尔滨工业大学交通科学与工程学院; 西交利物浦大学先进技术学院; 南安普顿大学电子与计算机科学学院; 苏州城市学院智能制造与智能交通学院; 青岛科技大学; 武汉大学数学与人工智能研究院; 复旦大学大数据研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有室外3D视觉 grounding 未充分利用多传感器互补物理属性的问题,提出首个多传感器数据集Talk2Sensors及TSFormer框架,在相关基准上实现了最优性能。

AI 中文摘要

作为具身智能的关键能力,3D视觉 grounding(3DVG)主要在采用RGB-D或点云输入的室内场景中被研究,而现有的室外扩展工作大多仅依赖单目图像。这两种设置都未能满足真实世界的室外感知需求,在真实室外场景中,异构传感器能捕获互补但特性不同的物理属性,如视觉纹理、3D几何和物体运动学,这些属性对于灵活且鲁棒的查询自适应 grounding 不可或缺,但至今未被充分利用。为弥合这一差距,我们推出了Talk2Sensors,这是首个基于相机、LiDAR和4D雷达构建的多传感器3D视觉 grounding 数据集,它包含8682条语言指令和20558个被引用物体,具有与传感器特定物理线索明确对齐的多样化提示。此外,我们提出了TSFormer,这是一种用于自动驾驶中语言引导的3D视觉 grounding 的统一Transformer框架。TSFormer采用由粗到细的属性感知融合策略:语言路由属性采样器首先通过用查询级语言线索调制传感器采样权重,执行粗粒度的文本条件特征检索;随后的稀疏保留模态仲裁器模块进行细粒度的模态仲裁和文本引导的细化,以确定精确的被引用空间位置。该设计能根据每个提示的语义需求动态路由外观、几何和运动线索,防止密集模态压倒稀疏但关键的传感器信号。大量实验表明,TSFormer在多个基准上达到了最先进的性能:它在Talk2Sensors上比最强基线提升了8.05 mAP,并且以53.05%的Acc@0.5迁移到单目Mono3DRefer基准。

英文摘要

As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inputs, while existing outdoor extensions largely rely on monocular images alone. Both settings fall short of real-world outdoor perception, where heterogeneous sensors capture complementary yet distinct physical properties, such as visual texture, 3D geometry, and object kinematics, that are indispensable for flexible and robust query-adaptive grounding but remain under-exploited. To bridge this gap, we introduce Talk2Sensors, the first multi-sensor 3D visual grounding dataset built upon camera, LiDAR, and 4D radar. It contains 8,682 language instructions and 20,558 referred objects, with diverse prompts explicitly aligned with sensor-specific physical cues. Furthermore, we propose TSFormer, a unified Transformer-based framework for language-guided 3D visual grounding in autonomous driving. TSFormer adopts a coarse-to-fine property-aware fusion strategy: the Language-Routed Property Sampler first performs coarse text-conditioned feature retrieval by modulating sensor sampling weights with query-level linguistic cues, while the subsequent Sparse-Preserving Modality Arbiter module conducts fine-grained modality arbitration and text-guided refinement to determine the precise referred spatial location. This design enables dynamic routing of appearance, geometry, and motion cues according to the semantic requirements of each prompt, preventing dense modalities from overwhelming sparse but critical sensor signals. Extensive experiments demonstrate that TSFormer achieves state-of-the-art performance across multiple benchmarks: it improves over the strongest baseline by 8.05 mAP on Talk2Sensors, and transfers to the monocular Mono3DRefer benchmark with 53.05\% Acc@0.5.

Comments14 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑