脉冲间隔变化对蝙蝠叫声深度学习分类的影响
Effects of interpulse-interval variation on deep-learning classification of bat vocalizations
浏览论文内容
中文总结 AI 辅助
本研究探讨脉冲间隔变化对蝙蝠叫声分类的影响,发现其对Transformer和CNN模型作用有限,但归一化条件结果难以直接迁移至自然录音。
中文摘要 AI 辅助
时间上下文可能有助于自动蝙蝠物种分类,但具体特征的贡献仍不清楚。我们研究了脉冲间隔(IPI)——连续叫声起始之间的时间——的变化是否提供物种判别信息,以及基于Transformer的模型是否比卷积神经网络对此信息更敏感。我们从欧洲蝙蝠录音中创建了两个匹配数据集:自然IPI条件保留原始叫声时序,以及归一化IPI条件,其中叫声起始以50毫秒间隔排列。EfficientNet-B0和PaSST在每个条件下进行微调和评估。在另一个实验中,每种架构分别在自然IPI和归一化IPI录音上训练,并在相同的自然IPI测试集上评估。最后,预训练分类器BatDetect2和BAT在两个条件下进行评估。条件内IPI归一化具有依赖模型的效果。PaSST在自然IPI(71±2.3%)和归一化IPI(70±6.3%)条件下的准确率差异很小,而EfficientNet的准确率从47±4.7%增加到57±3.9%。PaSST在两个条件下均超过EfficientNet。在跨条件评估中,在自然IPI录音上训练的模型在自然IPI测试集上优于在归一化IPI录音上训练的模型:EfficientNet的准确率从54%下降到50%,PaSST从65%下降到57%。BatDetect2和BAT在IPI条件之间差异很小。总体而言,我们发现有限的支持表明自然IPI变化对蝙蝠物种分类贡献显著,且Transformer模型比CNN模型更有效地利用它。然而,跨条件性能下降表明在归一化条件下获得的结果可能无法完全转移到自然录音中。
英文摘要
Temporal context may aid automated bat-species classification, but the contribution of specific features remains unclear. We investigated whether variation in the interpulse interval (IPI)-the time between consecutive call onsets-provides species-discriminative information and whether transformer-based models are more sensitive to this information than convolutional neural networks. We created two matched datasets from European bat recordings: a natural-IPI condition retaining the original call timing and a normalized-IPI condition in which call onsets were spaced at 50-ms intervals. EfficientNet-B0 and PaSST were fine-tuned and evaluated within each condition. In an additional experiment, each architecture was trained separately on natural-IPI and normalized-IPI recordings, and evaluated on the same natural-IPI test set. Finally, the pretrained classifiers BatDetect2 and BAT were evaluated on both conditions. Within-condition IPI normalization had model-dependent effects. PaSST accuracy differed little between the natural-IPI ($71 \pm 2.3\%$) and normalized-IPI ($70 \pm 6.3\%$) conditions, whereas EfficientNet accuracy increased from $47 \pm 4.7\%$ to $57 \pm 3.9\%$. PaSST exceeded EfficientNet under both conditions. In the cross-condition evaluation, models trained on natural-IPI recordings outperformed those trained on normalized-IPI recordings on the natural-IPI test set: accuracy decreased from 54% to 50% for EfficientNet and from 65% to 57% for PaSST. BatDetect2 and BAT differed little between IPI conditions. Overall, we found limited support for the hypotheses that natural IPI variation contributes substantially to bat-species classification and that it is used more effectively by transformer-based than CNN-based models. Nevertheless, the cross-condition performance decrease shows that results obtained under normalized conditions may not transfer fully to natural recordings.
发表机构
- Naturalis Biodiversity Centre(荷兰自然生物多样性中心)
- Tilburg University(蒂尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。