arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ECHOv2:用于异常声音检测的两级频带分割表示学习

ECHOv2: A Frequency-Structured Pre-trained Acoustic Representation Model with Cross-Band Modeling for Machine Anomalous Sound Detection

Yucong Zhang, Juan Liu, Ming Li

arXiv 2607.10596首次发表:更新:

发表机构

School of Computer Science, Wuhan University; School of Artificial Intelligence, Wuhan University; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(武汉大学计算机学院; 武汉大学人工智能学院; 香港中文大学(深圳)人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对机器异常声音检测中现有预训练音频主干不足的问题,提出ECHOv2模型,通过两级频带分割及自蒸馏策略学习表示,经实验验证其有效性,为异常声音检测表示学习提供强大框架并开源模型与基准。

AI 中文摘要

机器异常声音检测(ASD)需要强大的音频表示,以在有限监督下捕捉机器声音中的细微偏差。现有预训练音频主干不能完全捕捉机器声音的特定频率特征。为此提出ECHOv2,一种频带分割模型,学习局部带内表示以捕捉细粒度频谱模式,还采用两级自蒸馏策略及显式带间监督来建模跨频率依赖性。带间分支进行全局上下文对齐和掩码子带重建,引入多个摘要令牌进行结构化聚合。该设计使ECHOv2能稳健处理多种机器类型和嘈杂操作条件。建立统一ASD基准对预训练音频主干进行公平一致评估,消融研究证实了带内学习、带间监督和结构化聚合粒度对稳健ASD表示学习的有效性。这些发现表明结构化跨带建模为ASD表示学习提供了强大且适应性强的框架。模型和基准已完全开源。

英文摘要

Machine anomalous sound detection (ASD) is an important technology for industrial acoustic monitoring, where robust acoustic representation learning remains challenging due to limited anomalous samples and complex machine sound characteristics. Existing pre-trained acoustic representation models do not fully capture frequency-specific characteristics of machine sounds. To address this, we propose ECHOv2, a frequency-structured pre-trained acoustic representation model. The model learns localized intra-band representations to capture fine-grained spectral patterns while also incorporating a two-level self-distillation strategy with explicit inter-band supervision to model cross-frequency dependencies. The inter-band branch performs global context alignment and masked sub-band reconstruction, and multiple summary tokens are introduced for structured aggregation with controllable frequency granularity, enabling region-aware interaction across sub-bands during training. This design allows ECHOv2 to robustly handle diverse machine types and noisy operating conditions while maintaining stable representation quality. To enable fair and consistent evaluation of pre-trained audio backbones, we establish a unified ASD benchmark over DCASE 2020--2025 with two complementary protocols: embedding-based evaluation for frozen representation discriminability and adaptation-based evaluation for downstream transferability. Ablation studies confirm the effectiveness of intra-band learning, inter-band supervision, and structured aggregation granularity for robust ASD representation learning. These findings demonstrate that structured cross-band modeling improves acoustic representation learning for machine monitoring and provides an effective pre-trained representation model for industrial anomalous sound detection. The model and benchmark are publicly available to promote reproducible research.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑