Clean2FX:用于清音到效果音吉他音频转换的标签条件建模
Clean2FX: Label-conditioned modeling for clean-to-effect guitar audio transformations
浏览论文内容
中文总结 AI 辅助
研究电吉他音频清音到效果音转换,基于频谱图设置评估四种神经方法,包括两个变分自编码器和两个U-Net模型,U-Net模型表现更佳,不同效果改善情况有差异,还通过演示网站展示了模型应用效果。
中文摘要 AI 辅助
我们展示了Clean2FX,这是一项关于电吉他音频的标签条件清音到效果音转换的研究及演示。给定清音吉他输入和目标效果标签,任务是合成相应的效果信号并保留音乐内容。训练和评估对由EGFxSet真实单音录音构建,通过组合匹配的清音/效果和弦、旋律和混合时间线,便于跨效果进行对比。我们在基于频谱图的通用转换设置下评估了四种神经方法:两个变分自编码器和两个在处理线性或对数幅度表示上不同的U-Net模型。性能通过线性幅度频谱图MSE和Fréchet音频距离衡量。U-Net模型优于变分自编码器变体。每个效果的结果表明失真效果最易改善,而延迟和混响效果尽管频谱误差大幅降低,但FAD增益较弱。条件敏感性诊断表明最佳模型对目标标签有响应,而非退化为单一转换。我们的演示网站比较了应用于训练和验证数据之外的真实世界吉他演奏的两个模型,提供了实际清音到效果音行为的音频和频谱图示例。
英文摘要
We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio. Given a clean guitar input and a target effect label, the task is to synthesize the corresponding effected signal while preserving the musical content. Training and evaluation pairs are constructed from EGFxSet real, single tone recordings by assembling matched clean/effected chords, melodies, and mixed timelines. This allows for controlled comparison across effects. We evaluate four neural approaches under a common spectrogram-based transformation setting: two variational autoencoders and two U-Net models that differ in whether they operate on linear or log-magnitude representations. Performance is measured using linear-magnitude spectrogram MSE and Fréchet Audio Distance. The U-Net models outperform the variational autoencoder variants. Per-effect results show that distortion effects are most readily improved, whereas delay and reverb effects exhibit weaker FAD gains despite substantial spectral-error reductions. A conditioning-sensitivity diagnostic provides evidence that the best model responds to target labels rather than collapsing to a single transformation. Our demo website compares two models applied on real-world guitar performances outside training and validation data, providing audio and spectrogram examples of the practical clean-to-effect behavior.