用于自动驾驶汽车鲁棒交通标志感知的视觉-语言模型蒸馏
Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles
浏览论文内容
中文总结 AI 辅助
本研究提出LAMDA框架,通过冻结OpenCLIP文本编码器构建原型库监督视觉特征,在GTSRB和LISA数据集上,可同时提升TSR模型对三类物理攻击的鲁棒性且不降低干净准确率。
中文摘要 AI 辅助
基于深度神经网络的交通标志识别(TSR)模型在干净数据上表现出色,但易受物理可实现的对抗攻击影响,包括阴影扰动、自然光干扰和打印补丁。现有防御方法通常仅提升对某类攻击的鲁棒性,却会降低对其他攻击的性能,还可能降低干净准确率。我们提出LAMDA(Language-Anchored Model for Direction Alignment,语言锚定方向对齐模型),这是一种训练框架,无需使用对抗样本或增加推理时开销,即可将语言基础结构迁移至TSR模型。LAMDA利用冻结的OpenCLIP文本编码器,从视觉-语言模型(VLM)生成的标志描述和类别名称构建两个固定原型库,并在训练期间通过两种互补辅助损失监督视觉特征。推理时,适配器和原型库被丢弃,仅保留标准主干网络和分类器。在GTSRB和LISA数据集上,针对四个主干网络和三种物理攻击类型的评估显示,在被评估的十种方法中,LAMDA是唯一能在所有攻击-主干-数据集组合中持续提升鲁棒性的方法,在阴影攻击下提升达12.5个百分点,在自然光攻击下提升达13.2个百分点,同时在几乎所有情况下保留或提升干净准确率。
英文摘要
Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches. Existing defenses often improve robustness against one attack type while degrading performance on others, and can reduce clean accuracy. We propose LAMDA (Language-Anchored Model for Direction Alignment), a training framework that transfers language-grounded structure into TSR models without using adversarial examples or adding inference-time overhead. LAMDA builds two fixed prototype banks from VLM-generated sign descriptions and class names using a frozen OpenCLIP text encoder, and uses them to supervise visual features through two complementary auxiliary losses during training. At inference, the adapter and prototype banks are discarded, leaving a standard backbone and classifier. Evaluated on GTSRB and LISA across four backbones and three physical attack types, LAMDA is the only method among ten evaluated that consistently improves robustness across all attack-backbone-dataset combinations, with gains of up to +12.5 pp under shadow attacks and +13.2 pp under natural-light attacks, while preserving or improving clean accuracy in nearly all cases.
发表机构
- Clemson University(克莱姆森大学)
机构由 AI 辅助整理,请以论文原文为准。