arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向实车部署的鲁棒激光雷达语义分割:粗标签、恶劣环境与域偏移下的评估

Toward Robust LiDAR Semantic Segmentation for Real-World Deployment: Evaluation under Coarse Labels, Adverse Conditions, and Domain Shifts

Samir Abou Haidar, Alexandre Chariot, Mehdi Darouich, Cyril Joly, Jean-Emmanuel Deschaud

arXiv 2609.02830首次发表:更新:

发表机构

Mines Paris, PSL University; Paris-Saclay University; CEA(巴黎高等矿业学校(PSL大学); 巴黎萨克雷大学; 法国原子能和替代能源委员会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种结构化评估协议,从粗标签、恶劣环境干扰、域泛化及嵌入式平台推理速度维度评估LiDAR语义分割模型的部署就绪性,发现基准性能与实车部署需求存在差距,为相关评估提供参考。

AI 中文摘要

基于激光雷达(LiDAR)的语义分割是自动驾驶车辆与移动机器人的核心感知模块。尽管近期最先进的方法在标准基准上表现强劲,但现有评估协议仍聚焦于干净的单域设置与细粒度标签分类,基本未评估部署就绪性。实车系统必须处理关乎安全的标签语义、退化的感知条件及跨域变异性,但目前尚无统一协议同时覆盖这三个方面。本文提出一种结构化评估协议,沿三个互补维度评估LiDAR语义分割模型的部署就绪性:(i)契合自动驾驶安全优先级的粗标签评估,揭示标签粒度对不同方法的影响;(ii)八种模拟真实大气、几何及传感器退化的LiDAR干扰下的鲁棒性;(iii)无适配情况下跨数据集的域泛化。评估包含在嵌入式Jetson AGX Orin平台上测得的推理速度,直接反映部署约束。结果表明,细粒度基准排名未必反映安全相关性能,所有方法在干扰下均出现显著退化且鲁棒性具架构依赖性,当前域泛化仍不足以支持可靠部署。这些发现暴露了基准性能与部署就绪性间的具体差距,为更贴合实际的LiDAR语义分割评估提供参考协议。

英文摘要

LiDAR-based semantic segmentation is a core perception module for autonomous vehicles and mobile robots. Despite the strong performance of recent state-of-the-art methods on standard benchmarks, existing evaluation protocols remain focused on clean, single-domain settings and fine-grained label taxonomies, leaving deployment readiness largely unassessed. Real-world systems must handle safety-critical label semantics, degraded sensing conditions, and cross-domain variability, yet no unified protocol currently addresses all three aspects together. In this paper, we propose a structured evaluation protocol that assesses the deployment readiness of LiDAR semantic segmentation models along three complementary dimensions: (i) coarse-label evaluation aligned with autonomous driving safety priorities, revealing how label granularity affects different methods; (ii) robustness under eight types of LiDAR corruptions designed to emulate real-world atmospheric, geometric, and sensor degradations; and (iii) domain generalization across datasets without adaptation. The evaluation includes inference speed measured on an embedded Jetson AGX Orin platform, directly reflecting deployment constraints. Our results show that fine-grained benchmark rankings do not always reflect safety-relevant performance, that all methods experience substantial degradation under corruptions with architecture-dependent robustness characteristics, and that current domain generalization remains insufficient for reliable deployment. These findings expose concrete gaps between benchmark performance and deployment readiness, and provide a reference protocol for more practically grounded evaluation of LiDAR semantic segmentation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑