发表机构
Lawrence Technological University(劳伦斯理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自动驾驶中应急车辆检测的模态退化问题,提出AVNet多模态音视频Transformer,通过时间对齐交叉注意力融合和知识蒸馏,在AudioSet上达66.6%准确率,显著优于单模态。
AI 中文摘要
自动驾驶中的应急车辆检测是一项安全关键的感知任务,要求在多样且不利的真实世界条件下具备鲁棒性。现有方法依赖单一模态(音频或视频),当该模态退化时会导致系统性失败:基于麦克风的系统在嘈杂的城市环境中失效,而基于摄像头的系统在夜间或遮挡情况下失效。本报告提出AVNet,一种多模态音视频Transformer,利用音频和视频对应急车辆(救护车、消防车、警车)及道路背景进行分类,同时在推理时优雅地处理任一模态缺失的情况。AVNet引入三项关键贡献:(1)一个时间对齐的跨模态融合模块,在音频频谱图令牌和视频帧令牌之间执行秒级交叉注意力,利用它们精确的时间对应关系,无需任何学习到的对齐机制;(2)学习到的空嵌入,用于替代缺失的模态令牌,使单一统一模型无需重新训练即可在仅音频、仅视频或联合音视频模式下运行;(3)一种知识蒸馏训练策略,其中专家单模态教师模型通过软概率目标将类间暗知识迁移到多模态学生融合分支。在Google AudioSet数据集的281个片段上评估,AVNet在音视频模式下达到66.6%的总体准确率,比仅音频分支高出+10.4%,比仅视频分支高出+15.0%。最大的每类增益出现在最难类别“救护车”上,融合相比任一单模态分支单独高出+29.5%,表明两种模态提供了互补信息,且对齐的交叉注意力机制成功利用了这些信息。
英文摘要
Emergency vehicle detection in autonomous driving is a safety-critical perception task that demands robustness under diverse and adverse real-world conditions. Existing approaches rely on a single modality, either audio or video, which leads to systematic failure when that modality is degraded: microphone-based systems fail in noisy urban environments, and camera-based systems fail at night or under occlusion. This report presents AVNet, a multimodal audio-visual transformer that classifies emergency vehicles (ambulance, fire engine, police car) and road background using both audio and video, while gracefully handling the absence of either modality at inference time. AVNet introduces three key contributions: (1) a temporally aligned cross-modal fusion module that performs second-level cross-attention between audio spectrogram tokens and video frame tokens, exploiting their exact temporal correspondence without any learned alignment mechanism; (2) learned null embeddings that substitute for missing modality tokens, enabling a single unified model to operate in audio-only, video-only, or joint audio-visual mode without retraining; and (3) a knowledge distillation training strategy in which specialist unimodal teacher models transfer inter-class dark knowledge into the multimodal student fusion branch via soft probability targets. Evaluated on 281 clips from the Google AudioSet dataset, AVNet achieves 66.6% overall accuracy in audio-visual mode, outperforming the audio-only branch by +10.4% and the video-only branch by +15.0%. The largest per-class gain is observed for the hardest class, Ambulance, where fusion achieves +29.5% over either unimodal branch alone, demonstrating that the two modalities provide complementary information that the aligned cross attention mechanism successfully exploits.
Comments15 pages, 1 figure