arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17076cs.AIcs.SD

样本条件化表示选择用于音频少样本学习

Sample-Conditioned Representation Selection for Audio Few-Shot Learning

  • National Key Laboratory of Space Integrated Information System, Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所空间综合信息处理技术国家重点实验室)
  • School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)
  • School of Computer Science, Nanjing University(南京大学计算机学院)
  • School of Psychology, Shanghai Jiao Tong University(上海交通大学心理学院)

机构由 AI 辅助整理,请以论文原文为准。

Fengrui Liu, Ningxin Shen, Yi Li, Yiwei Fu, Feng Liu, Jiangmeng Li

AI总结:

针对少样本音频分类中前景-背景共现偏移问题,提出SAMPLESELECT方法,通过为每个输入预测固定预算特征掩码,在冻结编码器下提升OOD准确率,较全表示控制提高4.90-8.38个百分点。

AI中文摘要:

少样本音频分类器可能依赖前景-背景共现,当这些相关性发生变化时,分类器会失效。在SpurAudio上,由此产生的表示偏移是集中且类别相关的:对于ResNet12,前10%的通道解释了82.80%的零校正偏移贡献。我们提出SAMPLESELECT,它为每个输入独立预测固定预算的特征掩码,同时保持编码器和源分类器冻结。训练使用可微的Gumbel Top-k选择,结合前景分类和跨背景对比损失;推理使用确定性的Top-k掩码和仅支持集的线性适应。在ResNet12和Conv64的5-way 1-shot和5-shot评估中,SAMPLESELECT在比较方法中取得了最佳的OOD准确率,并将匹配的全表示控制提高了4.90-8.38个百分点。消融研究和表示分析进一步支持了所学习的选择机制。代码可在该https URL获取。

英文摘要:

Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propose SAMPLESELECT, which predicts a fixed-budget feature mask independently for each input while keeping the encoder and source classifier frozen. Training uses differentiable Gumbel Top-k selection with foreground classification and cross-background contrastive losses; inference uses deterministic Top-k masks and support-only linear adaptation. Across ResNet12 and Conv64 in 5-way 1-shot and 5-shot evaluation, SAMPLESELECT gives the best OOD accuracy among the compared methods and improves the matched full-representation control by 4.90-8.38 percentage points. Ablations and representation analyses further support the learned selection mechanism. Code is available at https://github.com/Cross-Innovation-Lab/SAMPLESELECT/

补充信息

↑