发表机构
City University of Hong Kong; Lingnan University(香港城市大学; 岭南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自动驾驶中感知模型因域偏移难以泛化的问题,提出LDE框架,利用协同感知生成高质量伪标签,通过特征共享、视野过滤和课程学习策略,在3D目标检测上优于现有方法。
AI 中文摘要
在自动驾驶中,感知模型常因域偏移而难以泛化到新环境。虽然无监督模型适配提供了一种无需繁重人工标注的可行解决方案,但现有仅依赖自车数据的方法往往导致较差的伪标签性能。为解决这一关键问题,我们提出LDE,即从分布式“眼睛”中学习,一种将协同感知(CP)转化为模型适配的高质量监督来源的新框架。这种伪标签方法对超参数不敏感且相对可靠,其假设是CP通常优于单智能体感知。然而,直接实施该方法会遇到以下问题:(1)在时间和带宽限制下共享丰富特征带来的通信瓶颈;(2)CP视角与学习者视野(FoV)之间的视角差异;(3)即使CP生成的标签也可能存在不可靠性。为解决这些问题,我们设计了一种面向适配的特征共享机制,选择性地传输对适配最关键的信息;一种FoV过滤方法,仔细消除不匹配的标签;以及一种课程学习策略,逐步利用伪标签。在3D目标检测任务上的大量实验表明,LDE始终优于预训练模型和最先进的无监督适配方法。
英文摘要
In autonomous driving, perception models often struggle to generalize to new environments due to domain shifts. While unsupervised model adaptation offers a feasible solution without labor-intensive manual labeling, existing methods that rely solely on the ego-vehicle's data often lead to inferior pseudo-labeling performance. To address this critical issue, we propose LDE, Learning from Distributed ``Eyes", a novel framework that transforms collaborative perception (CP) into a source of high-quality supervision for model adaptation. This pseudo-labeling approach is hyperparameter-insensitive and relatively reliable, assuming CP often outperforms single-agent's perception. However, naively implementing this approach encounters (1) the communication bottleneck of sharing rich features under time and bandwidth constraints, (2) the view discrepancy between the CP view and the learner's Field of View (FoV), and (3) the unreliability even in CP-generated labels. To address these issues, we design an adaptation-oriented feature sharing mechanism that selectively transmits the most critical information for adaptation, an FoV filtering method that meticulously eliminates mismatched labels, and a curriculum learning strategy to progressively exploit pseudo labels. Extensive experiments on 3D object detection tasks demonstrate that LDE consistently outperforms both the pre-trained models and state-of-the-art unsupervised adaptation methods.
Comments9 pages, 3 figures