arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向自动驾驶的结合语义属性的多模态交通标志检测

Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

Meda Lazar, Sourab Sridhar, Shashwata Gupta, Alexandra Tripcea, Varun Ravi, Senthil Yogamani

arXiv 2608.20874首次发表:更新:

发表机构

Arriver System Software S.r.l.; Qualcomm Auto Ltd Sweden Filial; Qualcomm Auto Ltd.; Qualcomm Technologies, Inc(Arriver系统软件有限责任公司; 高通汽车有限公司瑞典分公司; 高通汽车有限公司; 高通技术公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自动驾驶中交通标志检测的跨区域泛化差、远距离小目标检测弱、跟踪易受透视畸变影响的问题,提出多模态框架,结合LiDAR与相机,引入强度感知可变形融合模块、双运动模型跟踪器及语义属性分类流水线,在多国数据集上取得低漏检率,实现全球可泛化的商用级交通标志感知。

AI 中文摘要

可靠的交通标志检测是自动驾驶系统全球部署的前提,监管合规性与道路安全依赖于在不同区域、距离和天气条件下正确感知标志。尽管近期取得了进展,但基于视觉的方法仍面临三个根本局限:因各国间高度多样性导致的跨区域泛化能力差、远距离小目标检测性能下降(200米处交通标志仅占10×10像素)、车辆接近标志时发生的强非线性透视畸变下的时间跟踪脆弱。本文中,我们结合相机与激光雷达(Light Detection and Ranging, LiDAR)感知,解决鲁棒、远距离、区域无关的交通标志感知问题。我们提出一种多模态检测框架,其强度感知可变形融合模块将反射式LiDAR线索与相机特征对齐,以几何不变性而非区域特定视觉外观为基础锚定检测。我们还引入双运动模型跟踪器,明确考虑车辆接近时的非线性透视变换,相比线性运动假设大幅提升时间一致性。此外,我们开发语义属性分类流水线,估计遮挡程度、可读性、标志嵌入度及道路相关性,为下游规划提供可操作上下文。在我们的数据集(覆盖60余个国家、2500余小时驾驶数据)上开展的广泛评估显示,所提流水线在221068个评估序列中达到0.49%的目标漏检率(Object Miss Ratio, OMR),证明其在商用级自动驾驶系统中具备全球可泛化的交通标志感知能力。

英文摘要

Reliable traffic sign detection is a prerequisite for the global deployment of autonomous driving systems, where regulatory compliance and road safety depend on perceiving signs correctly across regions, ranges, and weather conditions. Despite recent progress, vision-based methods continue to face three fundamental limitations: poor cross-regional generalization due to high diversity across countries, degraded performance on small-object detection at long ranges (traffic signs occupy as little as $10{\times}10$ pixels at 200m), and fragile temporal tracking under the strongly non-linear perspective distortion that occurs as a vehicle approaches a sign. In this paper, we address the problem of robust, long-range, region-agnostic traffic sign perception by combining camera and Light Detection and Ranging (LiDAR) sensing. We present a multi-modal detection framework whose Intensity-Aware Deformable Fusion module aligns retro-reflective LiDAR cues with camera features, anchoring detection on geometric invariants rather than region-specific visual appearance. We further introduce a dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach, substantially improving temporal consistency over linear motion assumptions. Additionally, we develop a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedness, and road relevance, providing actionable context to downstream planning. Extensive evaluation on our dataset, spanning 60+ countries and 2,500+ hours of driving data, shows that the proposed pipeline achieves an Object Miss Ratio (OMR) of 0.49% across 221,068 evaluation sequences, demonstrating globally generalizable traffic sign perception in commercial-grade autonomous driving systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑