arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

仅基于相机的端到端驾驶的抗退化基准测试

A Degradation-Tolerance Benchmark for Camera-Only End-to-End Driving

Haohua Que, Handong Yao

arXiv 2608.29005首次发表:更新:

发表机构

College of Engineering, University of Georgia(佐治亚大学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员提出DriveDegrade基准测试,评估仅基于相机的端到端驾驶模型对各类图像退化的容忍度,发现不同退化对规划的影响差异显著,某视觉-语言-动作规划器被遮挡6个相机仅损失11.5%性能。

AI 中文摘要

仅基于相机的端到端(E2E)驾驶模型即将部署,而相机流可能会因模糊、噪声、弱光、天气、帧丢失和内存故障而退化。策略在驾驶性能失效前能承受多少退化尚不明确。现有的抗损坏基准测试针对的是检测或鸟瞰感知,而非驱动车辆的规划输出。我们提出了DriveDegrade,一个用于仅基于相机的端到端驾驶的图像退化容忍度基准测试。在图像加载器中实时注入16种损坏类别,每种类别分为5个严重程度,其中一种操作覆盖15种策略,我们在nuScenes和NAVSIM上评估开环规划,同时以CARLA作为闭环基准。研究发现:第一,轻度退化几乎不影响规划,导致规划失效的损坏类别在中等严重程度时有明确阈值;第二,脆弱性与损坏类型相关:模糊、JPEG和雨滴对规划的破坏最大,而天气和比特错误在很大程度上可被容忍;第三,平坦曲线具有歧义,因此我们将损坏分为降低图像质量和移除图像信息两类。读取相机的规划器在信息被移除时必然会损失精度,无论其在质量损失下的表现如何。在这两个维度上,规划器表现出明显差异,量化了自我状态捷径,且不会将无差异误判为鲁棒性。一个已发布的视觉-语言-动作规划器在两个维度上均呈平坦曲线,且其所有6个相机被完全遮挡仅导致11.5%的性能损失。

英文摘要

Camera-only end-to-end (E2E) driving models are nearing deployment, where the camera stream is degraded by blur, noise, low light, weather, frame loss, and memory faults. How much a policy tolerates before its driving breaks is unclear. Corruption-robustness benchmarks target detection or bird's-eye-view perception, not the planning output that drives the car. We present DriveDegrade, a benchmark for image-degradation tolerance in camera-only E2E driving. Sixteen corruption families at five severities are injected on the fly inside the image loader, one operator reaching fifteen policies, and we evaluate open-loop planning on nuScenes and NAVSIM plus a CARLA closed-loop anchor. First, mild degradation barely affects planning, and the families that break it have a clear threshold at mid severity. Second, fragility is corruption-dependent: blur, JPEG, and raindrop damage planning most, while weather and bit error are tolerated far into the range. Third, a flat curve is ambiguous, so we separate corruptions that degrade the image from those that remove it. A planner that reads its camera must lose accuracy when information is deleted, whatever it does under quality loss. On these two axes the planners separate sharply, quantifying the ego-status shortcut without mistaking indifference for robustness. A released vision-language-action planner is flat on both axes, and blinding all six of its cameras costs it only 11.5 percent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑