用于声学行人检测的跨模态知识蒸馏
Cross-Modal Knowledge Distillation for Acoustic Pedestrian Detection
浏览论文内容
中文总结 AI 辅助
针对仅音频行人检测中跨模态蒸馏贡献不明的问题,提出信任过滤蒸馏(TFD),选择性抑制教师监督,实验表明其主要带来操作点偏移而非判别力提升。
中文摘要 AI 辅助
仅依赖音频的行人检测对城市感知具有吸引力,但受限于较弱的声学线索。一种有吸引力的策略是跨模态知识蒸馏,即在训练期间由视频教师模型监督音频学生模型,从而使部署模型仅凭音频运行。然而,在该任务严重的类别不平衡和视频-音频模态差异较大的情况下,这种蒸馏的贡献尚不明确。我们引入了信任过滤蒸馏(TFD),它选择性地抑制教师模型对行人样本的监督,并将其logit公式解释为在共享温度下的条件标签平滑。在ASPED数据集上进行的五折跨会话验证中,跨十种蒸馏配置,大多数方法在宏平均准确率上取得了适度提升,同时PR-AUC变化较小。主要效果是提高了无行人准确率,但以降低行人召回率为代价。将TFD添加到logit蒸馏中强化了这一权衡,但未改善平均PR-AUC。这些发现通过区分操作点偏移与判别力提升,阐明了在不平衡声学检测中选择性跨模态监督的益处与局限性。
英文摘要
Audio-only pedestrian detection is attractive for urban sensing but limited by weak acoustic cues. An appealing strategy is cross-modal knowledge distillation, in which a video teacher supervises the audio student during training so that the deployed model runs on audio alone. Under this task's severe class imbalance and wide video-audio modality gap, however, what such distillation contributes is unclear. We introduce Trust-Filtered Distillation (TFD), which selectively suppresses teacher supervision on pedestrian samples, and interpret its logit formulation as conditional label smoothing under a shared temperature. Across ten distillation configurations under five-fold cross-session validation on ASPED, most methods yield modest gains in macro accuracy accompanied by small changes in PR-AUC. The main effect is higher no-pedestrian accuracy at the cost of lower pedestrian recall. Adding TFD to logit distillation strengthens this trade-off without improving mean PR-AUC. These findings clarify the benefits and limitations of selective cross-modal supervision by distinguishing operating-point shifts from discrimination gains in imbalanced acoustic detection.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。