SPECTRA:用于全少样本类增量音频分类的子空间保留嵌入校准、迁移与重放
SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification
浏览论文内容
中文总结 AI 辅助
针对全少样本类增量音频分类的性能下降问题,提出含嵌入校准适配器、子空间特征重放、直推式最优传输优化的SPECTRA框架,在三类基准上优于现有方法。
中文摘要 AI 辅助
全少样本类增量音频分类(FFCAC)要求在每个会话中仅用少量标注示例识别新声音类别,同时不遗忘已学类别且无需大型基础数据集。现有方法通常冻结预训练的音频-语言编码器,用点原型进行分类,但因通用特征表示,在各会话中会出现显著性能下降。我们提出SPECTRA,这是一个基于冻结编码器的框架,新增三个组件:(i)轻量可训练适配器,将通用嵌入校准至对应任务;(ii)子空间特征重放,一种无样本的抗遗忘方案,通过从旧类别存储特征的低秩子空间采样来重放旧类别;(iii)测试时原型的直推式最优传输优化。我们的核心发现是,重放的子空间结构可减轻遗忘,且优于等方差的朴素高斯重放。在三个FFCAC基准(NSynth-100、FSC-89、LS-100)上,SPECTRA提升了平均准确率并降低了遗忘,超过当前最先进方法, ablation实验从统计上验证了每个组件的有效性。
英文摘要
Fully few-shot class-incremental audio classification (FFCAC) requires recognizing new sound classes from only a handful of labeled examples per session, without forgetting previously learned classes and without any large base dataset. Existing methods typically freeze a pre-trained audio--language encoder and classify with point prototypes, but they suffer from significant performance degradation throughout the sessions due to generic feature representations. We propose SPECTRA, a framework built on a frozen encoder which adds three components. (i) a lightweight trainable adapter that calibrates the generic embeddings to the task; (ii) subspace feature replay, an exemplar-free anti-forgetting scheme that replays old classes by sampling from the low-rank subspace of their stored features; and (iii) a transductive optimal-transport refinement of prototypes at test time. Our central finding is that the subspace structure of the replay diminishes forgetting and outperforms naive Gaussian replay of equal variance. On three FFCAC benchmarks (NSynth-100, FSC-89, LS-100), SPECTRA improves average accuracy and reduces forgetting over current state-of-the-art methods, and our ablations statistically validate each component.
发表机构
- University of Haifa(海法大学)
- University of Stuttgart(斯图加特大学)
机构由 AI 辅助整理,请以论文原文为准。