arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22830cs.CE

旋翼自噪声下经真实数据验证的无人机听觉用于低误报人体检测

Real-Validated UAV Audition Under Rotor Ego-Noise for Low-False-Alarm Human Detection

  • Faculty of Applied Sciences, Macao Polytechnic University(澳门理工学院应用科学学院)
  • College of Animal Science and Technology, Zhongkai University of Agriculture and Engineering(仲恺农业工程学院动物科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Junhao Wei, Haochen Li, Dexing Yao, Yanxiao Li, Yifu Zhao, Baili Lu, Zhenhong Peng, Ngai Cheong, Xu Yang, Yapeng Wang

AI总结:

针对无人机旋翼自噪声下人体声学检测,提出经真实数据验证的基准与评估协议,证明合成准确率不能可靠预测真实迁移性能,需分组级真实验证。

AI中文摘要:

利用无人机搭载的麦克风检测人类声学线索可支持声学搜索与救援,但旋翼自噪声常常在极低信噪比下掩盖语音、哭喊、咳嗽及其他人类声音。我们在此真实运行约束下研究无人机人类可闻存在性检测。模型在基于公共音频构建的可复现合成混合流水线上训练,但在真实DroneAudioSet录音上通过元数据定义的听觉滤波器、按录音分组的开发/测试划分及分组自助置信区间进行选择和评估。结果表明,合成准确率是真实无人机迁移的弱且非单调的代理指标:从头训练的SE-ResNet在合成混合数据上看似有竞争力,但在真实自噪声下崩溃,而冻结的音频基础模型和轻量适配器迁移更可靠。我们进一步评估了带有旋翼感知条件化和域正则化的BEATs适配器系列。真实开发集选出的EgoRAP-DA配置在候选方案中实现了最佳锁定测试低误报召回率,但其相对于普通适配器的优势在配对分组自助检验下不具有统计显著性。因此,主要贡献是一个经真实数据验证的基准和评估协议,表明无人机听觉的诚实进步需要真实、分组级别的验证,而非仅依赖合成分数。

英文摘要:

Detecting human acoustic cues from UAV-mounted microphones could support acoustic search and rescue, but rotor ego-noise often masks speech, cries, coughs, and other human sounds at extremely low SNRs. We study UAV human-audible-presence detection under this real operating constraint. Models are trained on a reproducible synthetic mixture pipeline built from public audio, but selected and evaluated on real DroneAudioSet recordings using a metadata-defined audibility filter, recording-grouped Dev/Test splits, and group-bootstrap confidence intervals. Our results show that synthetic accuracy is a weak and non-monotonic proxy for real UAV transfer: a from-scratch SE-ResNet appears competitive on synthetic mixtures but collapses on real ego-noise, while frozen audio foundation models and lightweight adapters transfer more reliably. We further evaluate a BEATs adapter family with rotor-aware conditioning and domain regularization. The Real-Dev-selected EgoRAP-DA configuration achieves the best locked-test low-false-alarm recall among the candidates, but its advantage over a vanilla adapter is not statistically significant under paired group bootstrap. The main contribution is therefore a real-validated benchmark and evaluation protocol showing that honest progress in UAV audition requires real, group-level validation rather than synthetic scores alone.

↑