从掩蔽到合并:重新思考SpecAugment以实现高效的音频频谱图Transformer
From Masking to Merging: Rethinking SpecAugment for Efficient Audio Spectrogram Transformer
浏览论文内容
中文总结 AI 辅助
本文提出SpecAugment-Patch Merging方法,通过掩蔽并合并补丁对减少AST处理的令牌数,在几乎不损失准确率(AudioSet mAP 34.07→34.08)的情况下提升吞吐量13.9%,实现高效训练。
中文摘要 AI 辅助
本文提出SpecAugment-Patch Merging,一种简单而有效的方法来加速音频频谱图Transformer(AST)的训练。我们首先应用SpecAugment在补丁级别掩蔽输入频谱图,并在添加位置嵌入后,该方法选择r对掩蔽补丁并将其合并,从而减少Transformer处理的令牌数量。将合并对的数量r从0增加到100,AudioSet上的mAP几乎保持不变(从34.07到34.08),而吞吐量从43.3增加到49.3样本/秒,相对提升了13.9%。类似模式出现在ESC-50和Speech Commands V2上,吞吐量稳步提高而准确率仅有微小变化,表明这种合并方法能以极小的性能损失实现更快的训练。
英文摘要
This paper proposes SpecAugment-Patch Merging, a simple yet effective method to accelerate Audio Spectrogram Transformer (AST) training. We first apply SpecAugment to mask input spectrograms at the patch level, and after positional embeddings are added, the method selects r pairs of masked patches and merges them, reducing the number of tokens processed by the Transformer. Increasing the number of merged pairs r from 0 to 100 keeps mAP on AudioSet nearly unchanged (34.07 to 34.08) while throughput increases from 43.3 to 49.3 samples/sec, which is a relatively 13.9% improvement. Similar patterns appear on ESC-50 and Speech Commands V2, where throughput steadily improves with only minor accuracy changes, demonstrating that this merging approach provides faster training with minimal performance loss.
发表机构
- Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。