基于证据融合与冲突折扣信念聚合的不确定性感知多模态反无人机检测
Uncertainty-Aware Multimodal Anti-UAV Detection via Evidential Fusion and Conflict-Discounted Belief Aggregation
浏览论文内容
中文总结 AI 辅助
针对现有多模态反无人机检测未建模不确定性的问题,本文扩展EDTC提出DBF方法实现多模态融合,在Anti-UAV基准上达到0.670准确率、至少38 FPS,校准良好但定位失败检测能力弱于空间方差。
中文摘要 AI 辅助
反无人机感知系统在传感器流因遮挡、快速运动或模态特定故障而退化时,必须保持可靠性。现有多模态反无人机系统确定性地融合RGB和热红外流,未对预测不确定性建模,且在流不一致时无法表达疑虑。证据深度学习(EDL)可在单次前向传播中生成校准后的每类不确定性,EDTC已将其用于仅热红外感知,但跨模态证据融合仍未解决。本文通过折扣信念融合(DBF)将EDTC扩展至多模态RGB-热红外感知,该方法在聚合流意见前将跨模态冲突转换为不确定性质量,通过选择不确定性较低的模态确定边界框。在Anti-UAV基准上,多模态融合在实时速度(至少38 FPS)下始终优于任一单流(测试准确率0.670,对比红外流0.604、RGB流0.598)。然而,DBF与未折扣平均法在经验上无差异:该以存在性为主的基准上跨模态冲突接近零,使折扣步骤无效。融合后的不确定性校准良好(预期校准误差ECE为0.057),但作为定位失败检测器的表现弱于空间方差(AUROC为0.626,对比0.739)。该无效结果具有结构性:基准的近乎普遍存在性与空漏检编码共同抑制了跨模态冲突,此诊断明确了感知冲突感知融合可提供可测量收益的场景。
英文摘要
Anti-UAV perception systems must remain reliable when sensor streams degrade under occlusion, fast motion, or modality-specific failure. Existing multimodal anti-UAV systems fuse RGB and thermal streams deterministically, without modeling predictive uncertainty, and cannot express doubt when streams disagree. Evidential Deep Learning (EDL) produces calibrated per-class uncertainty in a single forward pass. EDTC already exploits this for thermal-only perception, yet cross-modal evidential fusion remains unaddressed. This paper extends EDTC to multimodal RGB-Thermal perception via Discounted Belief Fusion (DBF), which converts inter-modal conflict into uncertainty mass before aggregating stream opinions. Bounding boxes are resolved by selecting the lower-uncertainty modality. On the Anti-UAV benchmark, multimodal fusion consistently outperforms either single stream (test Acc 0.670 vs. 0.604 IR, 0.598 RGB) at real-time speed (at least 38 FPS). However, DBF is empirically indistinguishable from undiscounted averaging: near-zero inter-modal conflict on this presence-dominated benchmark leaves the discounting step inert. The fused uncertainty is well-calibrated (ECE 0.057) yet expectedly a weaker localization failure detector than spatial variance (AUROC 0.626 vs. 0.739). The null result is structural: the benchmark's near-universal presence and vacuous miss-encoding jointly suppress inter-modal conflict, a diagnosis that delimits where conflict-aware fusion provides measurable benefit.
发表机构
- University of Amsterdam(阿姆斯特丹大学)
- SUNY Empire State University(纽约州立大学帝国州立大学)
机构由 AI 辅助整理,请以论文原文为准。