arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向增量音频分类的领域特定专家参数隔离方法

Parameter isolation with domain-specific experts for incremental audio classification

Jongyeon Park, Do-Hyeon Lim, Sang-won Park, Hong Kook Kim, Kyungdeuk Ko, Hyeongcheol Geum, Jeong Eun Lim

arXiv 2609.14730首次发表:更新:

发表机构

Hanwha Vision Co. Ltd.(韩华视觉有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种领域特定参数隔离架构,通过全阶循环更新、无数据生成重放和跨领域特征生成,在音频分类的域增量学习中缓解灾难性遗忘,并在DCASE 2026挑战赛任务7上分别将微观和宏观准确率提升至78.4%和78.9%。

AI 中文摘要

为了在流式数据预测和传感控制等时变环境中成功部署模型,域增量学习(DIL)引起了关注,因为它旨在使先前训练的模型适应新到达的领域,同时在不访问早期领域数据的情况下保留其知识。跨领域的增量学习可视为一种循环更新,其中当前模型通过对先前领域继承的模型进行更新而获得。依赖域不变特征学习和权重正则化的传统DIL方法会逐渐覆盖或约束先前领域学到的参数,导致灾难性遗忘。相反,本文提出了一种新的领域特定参数隔离架构,保留所有过去的领域。所提出的架构通过全阶循环更新缓解灾难性遗忘,利用以所有先前冻结模型为条件的领域特定数据构建新的专家。为实现这一点,我们引入了无数据生成重放以重建先前领域数据,以及跨领域特征生成以恢复早期领域样本中缺失的后续专家特征。最后,我们将所提出的模型架构应用于DCASE 2026挑战赛任务7中定义的音频分类的域无关增量学习。最终,我们分别实现了78.4%和78.9%的微观和宏观准确率,相比挑战赛基线分别提高了33和25个百分点。进行了消融研究,以检验每个处理组件在分类准确率方面的有效性。

英文摘要

To successfully deploy a model in time-varying environments such as streaming data prediction and sensing control, domain-incremental learning (DIL) has attracted attention since it aims to adapt a previously trained model to newly arriving domains, while reserving knowledge from earlier domains without accessing their data. Incremental learning across domains can be regarded as a recurrent update, in which the current model is obtained by updating the model carried over from previous domains. Conventional DIL approaches that rely on domain-invariant feature learning and weight regularization gradually overwrite or constrain parameters learned in previous domains, leading to catastrophic forgetting. Instead, this paper proposes a new domain-specific parameter-isolation architecture that retains all past domains. The proposed architecture mitigates catastrophic forgetting through a full-order recurrent update, constructing a new expert using domain-specific data conditioned on all previously frozen models. To achieve this, we incorporate data-free generative replay to reconstruct previous-domain data and cross-domain feature generation to recover later expert features missing from earlier domain samples. Finally, we apply the proposed model architecture to domain-agnostic incremental learning for audio classification, as defined in the DCASE 2026 Challenge Task 7. Consequently, we achieve micro and macro accuracies of 78.4% and 78.9%, respectively, representing increases of 33 and 25 percentage points over the Challenge baseline. Ablation studies are conducted to examine the effectiveness of each processing component in terms of classification accuracy.

Comments5 pages, 3 figures, 3 tables. Accepted to the Detection and Classification of Acoustic Scenes and Events (DCASE) Workshop 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑