AI 中文总结
研究如何学习混音风格表示,提出StemFX框架,通过自回归预测源分离音轨上的FX链,Transformer解码器与多频段CNN编码器协同工作,经大规模训练后在混音风格检索和转换任务中表现出色,远超基线模型和迭代优化。
AI 中文摘要
音频混音风格涵盖混音工程师的艺术和技术决策,包括电平平衡、空间化以及每个音轨上音频效果(FX)的选择、排序和参数设置。FX链是这种风格的关键决定因素,但现有建模方法有限。我们提出了StemFX框架,通过自回归预测源分离音轨上的可变长度FX链来学习混音风格表示。Transformer解码器自回归预测令牌化的FX链,带分割多频段CNN编码器结合FiLM条件捕捉每个音轨的频谱结构。为实现大规模配对训练,我们通过源分离从约105K首歌曲中提取伪音轨并使用MultiAFx进行增强。在混音风格检索评估中,StemFX在所有测试链长度上均优于所有基线模型。在配对混音风格转换中,StemFX实现了最佳频谱保真度和最高听众偏好,比迭代优化快4000倍以上。
英文摘要
Audio mixing style encompasses the artistic and technical decisions a mix engineer makes, including level balancing, spatialization, and the choice, ordering, and parameterization of audio effects (FX) on each stem. FX chains are a key determinant of this style, yet existing approaches to modeling them remain limited. Some operate on stereo mixtures without explicit per-stem FX chain modeling, others fix the number or type of effects per track, and many require differentiable effect implementations or scarce multitrack datasets. We present StemFX, a framework that learns mixing style representations by autoregressively predicting variable-length FX chains on source-separated stems. A Transformer decoder predicts tokenized FX chains autoregressively, while a band-split multi-band CNN encoder with FiLM conditioning captures per-stem spectral structure. To enable large-scale paired training, we extract pseudo-stems from about 105K songs via source separation and augment them using MultiAFx, a toolkit unifying 85 audio effects from 7 Python libraries. Evaluated on mixing style retrieval, StemFX outperforms all baseline models across all tested chain lengths. On paired mixing style transfer, StemFX achieves the best spectral fidelity and the highest listener preference, over 4000 times faster than iterative optimization.
CommentsAccepted to ISMIR 2026. 8 pages, 4 figures