LAST:循环音频频谱图变换器
LAST: Looped Audio Spectrogram Transformer
浏览论文内容
中文总结 AI 辅助
针对变换器深度增加带来的计算开销问题,提出循环音频频谱图变换器(LAST),通过复用模块仅细化类别令牌,在AudioSet上以更少参数和计算量超越十二层顺序变换器,并提升鲁棒性与泛化能力。
中文摘要 AI 辅助
增加变换器模型的深度可以提高识别性能,但代价相当高昂。每增加一层就需要更多参数,使得过程在计算上效率低下。我们探究是否可以将额外的处理专注于整合已计算出的特征。循环音频频谱图变换器(LAST)首先处理所有令牌,然后复用相同的模块,在固定的音频特征上仅细化类别令牌,从而使后续的传递变得廉价。在AudioSet上,十次传递的LAST实现了0.345的平均精度均值,相对优于十二层顺序变换器2.1%,同时参数减少了49.4%,乘法累加操作减少了42%,实测吞吐量提高了9.8%。在分别训练的模型中,将传递次数从两次增加到十次,在仅增加1.2%计算量的情况下提高了准确性。进一步的评估显示,对时间掩蔽和各种其他音频增强的鲁棒性有所提高,在音乐、环境和事件声音的分类任务上具有更好的泛化能力。
英文摘要
Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- University of California San Diego(加利福尼亚大学圣迭戈分校)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。