发表机构
University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多模态反无人机检测,通过在三个基准上的消融实验,发现证据深度学习的训练目标可提升检测准确率与错误排序能力,而登普斯特-谢弗融合、时序门控等组件未达预期效果。
AI 中文摘要
反无人机系统日益融合多种传感器,但其检测头未提供各模态的可靠性信号。本研究通过在三个基准上进行受控消融实验,评估证据深度学习(EDL)头、登普斯特-谢弗(DS)证据融合及不确定性驱动的时序传感器门控是否能提升反无人机检测性能,三个基准分别为热跟踪(AntiUAV600)、RGB-音频-RF分类(TRIDENT)及RGB-IR跟踪(MM-UAV)。EDL训练目标较重新训练的sigmoid基线提升了准确率(E1中准确率提升5.9个百分点,弃权(不执行)的跟踪器数量增至三倍;E2中分类准确率提升4.8个百分点,经剪辑聚类自举验证,p值为0.011),且能更优地对分类错误进行排序(熵UAUC约为0.94,而sigmoid基线为0.51)。其余组件未支持其各自假设:DS融合未优于简单概率平均;Dirichlet vacuity除预测熵外未增加排序能力,且在检测层面出现反转,极端背景不平衡使其编码类成员关系而非错误可能性,熵和sigmoid置信度也存在该问题;时序门控仅在几乎不激活时保留准确率,在共享骨干硬件上未实现延迟节省。因此,证据学习的益处主要源于其训练目标而非不确定性估计;裁剪级控制进一步将检测层面的故障定位至锚点级评估,而非学习到的表示。
英文摘要
Anti-UAV systems increasingly fuse multiple sensors, yet their detection heads provide no per-modality reliability signal. This study evaluates whether evidential deep learning (EDL) heads, Dempster-Shafer (DS) evidence fusion, and uncertainty-driven temporal sensor gating improve anti-UAV detection through a controlled ablation on three benchmarks: thermal tracking (AntiUAV600), RGB-audio-RF classification (TRIDENT), and RGB-IR tracking (MM-UAV). The EDL training objective improves accuracy over retrained sigmoid baselines (+5.9 percentage points in accuracy and a tripled tracker-on-absent rate in E1; +4.8 percentage points in classification accuracy in E2, surviving a clip-clustered bootstrap, p = 0.011) and ranks classification errors substantially better (entropy UAUC approximately 0.94 vs. 0.51). The remaining components do not support their respective hypotheses. DS fusion does not outperform simple probability averaging. Dirichlet vacuity adds no ranking power beyond predictive entropy and inverts at the detection level, where extreme background imbalance causes it to encode class membership rather than error likelihood, a failure also observed for entropy and sigmoid confidence. Temporal gating preserves accuracy only when nearly inactive and yields no realised latency saving on shared-backbone hardware. The benefit of evidential learning therefore arises primarily from its training objective rather than its uncertainty estimate; a crop-level control further localises the detection-level breakdown to anchor-level evaluation rather than the learned representation.