重新思考程序化音频预训练:源规模与目标适配
Rethinking Procedural Audio Pre-training: Source Scaling and Objective Adaptation
浏览论文内容
中文总结 AI 辅助
本文重新思考程序化音频预训练,区分公式类覆盖度与类内渲染多样性两种规模,发现其益处依赖学习范式与任务,并揭示程序化音频偏好低掩码率,提出源感知预训练方法。
中文摘要 AI 辅助
程序化音频已成为可迁移音频表示学习的可行来源,但其设计原则仍未被充分探索。本文重新审视两个问题:程序化源应如何扩展规模,以及在自然音频上开发的训练选择是否应原封不动地迁移到程序化音频。通过受控源,我们将规模区分为公式类覆盖度C和类内渲染多样性D。基于FDSL和AudioMAE的实验表明,这两种规模形式提供不同的益处,且依赖于学习范式和下游任务。一项匹配的AudioMAE研究进一步表明,程序化音频偏好低掩码率(10%--25%),而AudioSet-28K偏好50%--75%。共享码本分析揭示了程序化音频中较低的补丁多样性和较强的时间可预测性。这些结果促使我们提出源感知的程序化预训练,其中源扩展和学习配置被共同考虑。代码可在提供的链接获取。
英文摘要
Procedural audio has emerged as a viable source for transferable audio representation learning, but its design principles remain unclear.We revisit two questions: how a procedural source should be scaled, and whether training choices developed on natural audio should transfer unchanged to procedural data.Using a controlled source, we separate scale into formula-class coverage C and within-class rendering diversity I.Experiments with FDSL and AudioMAE show that these two forms of scale provide different benefits and depend on the learning formulation and downstream task. A matched AudioMAE study further shows that procedural audio favors low mask ratios (10%--25%), whereas AudioSet-28K favors 50%--75%. Shared-codebook analysis reveals lower patch diversity and stronger temporal predictability in procedural audio. These results motivate source-aware procedural pre-training, where source scaling and learning configuration are considered jointly.Code is available at https://github.com/Cross-Innovation-Lab/Formula-Bank.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- East China Normal University(华东师范大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。