发表机构
Federal University of Santa Catarina (UFSC)(圣卡塔琳娜联邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在CARLA中评估跨视角变换器预测鸟瞰图作为行为克隆策略输入,提出六通道表示和KDE加权,发现全局IoU不能预测驾驶性能,路线通道是导航关键瓶颈。
AI 中文摘要
在自动驾驶中,鸟瞰图(BEV)表示提供了车辆周围环境的结构化、自上而下的抽象,并已成为行为克隆(BC)策略的关键输入模态。虽然仿真中可轻松获得真实鸟瞰图地图,但实际部署需要将其替换为相机预测的地图——这种替换引入了感知误差,其对闭环驾驶性能的下游影响尚不明确。在本工作中,我们研究了在CARLA仿真器中,将跨视角变换器(CVT)预测的鸟瞰图地图作为BC智能体的直接策略输入。我们提出了一种六通道鸟瞰图表示,涵盖路面、规划路线、车道边界、车辆、行人和交通灯,并引入了一种核密度估计(KDE)加权方案,该方案重新平衡分割损失,使其更关注欠表示的驾驶操作,如弯道和交叉口。在CARLA两个城镇上的闭环评估表明,尽管KDE加权模型未达到最高的总体IoU,但它是唯一一个在无违规情况下完成完整回合的预测鸟瞰图智能体。这一差异揭示了全局分割指标是驾驶性能的较差代理指标:决定导航成功的是几何关键位置的预测质量,而路线通道成为在预测鸟瞰图输入下可靠智能体导航的主要瓶颈。
英文摘要
In autonomous driving, Bird's-Eye View (BEV) representations provide a structured, top-down abstraction of the vehicle's surroundings and have become a key input modality for Behavioral Cloning (BC) policies. While ground-truth BEV maps are readily available in simulation, real-world deployment requires replacing them with camera-predicted counterparts - a substitution that introduces perceptual errors whose downstream impact on closed-loop driving performance is not well understood. In this work, we investigate the use of Cross-View Transformer (CVT)-predicted BEV maps as direct policy inputs for a BC agent in the CARLA simulator. We propose a six-channel BEV representation covering road surface, planned route, lane boundaries, vehicles, pedestrians, and traffic lights, and introduce a Kernel Density Estimation (KDE) weighting scheme that rebalances the segmentation loss towards underrepresented driving maneuvers such as curves and intersections. Closed-loop evaluation across two CARLA towns shows that the KDE-weighted model is the only predicted-BEV agent to complete a full episode without infractions, despite not achieving the highest aggregate IoU. This discrepancy reveals that global segmentation metrics are poor proxies for driving performance: what determines navigation success is prediction quality at geometrically critical locations, and the route channel emerges as the primary bottleneck for reliable agent navigation under predicted BEV inputs.