难以承受之重:无人机音频分类的模型与方法缩放
The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification
浏览论文内容
中文总结 AI 辅助
针对无人机音频分类数据稀缺问题,系统比较多种模型与微调方法,发现轻量CNN结合选择性批归一化微调(更新<0.5%参数)达97.65%准确率,缩放方法优于缩放模型。
中文摘要 AI 辅助
随着无人机在消费和国防领域日益普及,从有限的、特定模态的数据中可靠地对它们进行分类是一项紧迫的挑战。主流方法是使用大型预训练网络并在任务数据上进行全量微调,这带来了沉重的计算和内存负担,在资源受限的无人机部署场景中难以承受,因为边缘推理和对新兴平台的快速重训练都是必需的。本文系统地跨越模型架构和微调方法两个维度对无人机音频分类进行缩放研究,探讨何时这种负担是合理的,何时更轻量的替代方案更优。使用一个包含31种无人机类别、共3,100条音频片段的自建数据集,我们在全量微调、仅分类器微调以及四种参数高效微调(PEFT)方法(SSF、IA3、OFT和选择性批归一化微调)下,评估了Transformer(ViT、AST)和卷积(自定义CNN、ResNet-18/152、MobileNet-V3-S/L、EfficientNet-B0/B7)骨干网络。所有配置均通过5折交叉验证,评估指标包括准确率、训练时间、可训练参数占比和推理时内存占用。对EfficientNet-B7进行选择性批归一化微调并采用三倍数据增强,在更新不到0.5%模型参数的情况下,取得了最高的验证准确率(97.65% ± 0.30)。在整个扫描中,轻量级CNN在准确率和效率上均持续优于Transformer。在数据稀缺的无人机音频分类任务中,缩放方法优于缩放模型。
英文摘要
As unmanned aerial vehicles (UAVs) become increasingly prevalent in consumer and defense settings, classifying them reliably from limited, modality-specific data is an urgent challenge. The dominant approach, large pretrained networks fully fine-tuned on task data, carries a substantial computational and memory weight that is hard to bear in resource-constrained UAV deployments, where edge inference and rapid retraining for emerging platforms are both required. This paper systematically scales across both model architectures and fine-tuning methods for UAV audio classification, asking when that weight is justified and when lighter alternatives prevail. Using a custom dataset of 3,100 audio clips spanning 31 drone classes, we evaluate transformer (ViT, AST) and convolutional (custom CNN, ResNet-18/152, MobileNet-V3-S/L, EfficientNet-B0/B7) backbones under full fine-tuning, classifier-only fine-tuning, and four parameter-efficient fine-tuning (PEFT) methods: SSF, IA3, OFT, and selective batch-norm tuning. All configurations are evaluated with 5-fold cross-validation across accuracy, training time, trainable-parameter share, and inference-time memory footprint. Selective batch-norm fine-tuning of EfficientNet-B7 with three-fold augmentations achieves the highest validation accuracy (97.65% +- 0.30) while updating under 0.5% of model parameters. Across the sweep, lightweight CNNs consistently outperform transformers on both accuracy and efficiency. For UAV audio classification under data scarcity, scaling the method outperforms scaling the model.
发表机构
- College of Charleston(查尔斯顿学院)
机构由 AI 辅助整理,请以论文原文为准。