EG-VAE:用于电吉他音色迁移与去除的统一框架
EG-VAE: A Unified Framework for Electric Guitar Tone Transfer and Removal
AI总结:
本文提出EG-VAE统一框架,通过变分自编码器解耦内容与音色表示,联合完成电吉他音色迁移与去除任务,经实验验证其性能优于任务特定基线方法。
AI中文摘要:
电吉他音色迁移(EGTT)与音色去除(EGTR)是吉他音色建模中的两项基础任务:EGTT将录音的音色替换为参考音色,而EGTR则从经过处理的湿信号录音中恢复干直接输入(DI)信号。尽管二者高度相关,现有研究却分别处理这两项任务,且均未取得令人满意的结果。本文提出EG-VAE,这是一种统一框架,通过变分自编码器从湿信号录音中解耦帧级内容与全局音色表示,联合建模EGTT与EGTR。EGTT通过将源内容与参考音色重组实现,EGTR则通过一种新颖的音色掩码目标实现,该目标在训练期间强制内容-音色解耦,并在推理阶段完成去除操作。为提升对训练中未见音色的迁移能力,第二训练阶段通过变分采样与音频效果增强塑造平滑音色空间。客观与主观评估的实验结果表明,EG-VAE在迁移与去除任务上均优于任务特定的基线方法,演示可在该https URL获取。
英文摘要:
Electric guitar tone transfer (EGTT) and tone removal (EGTR) are two fundamental tasks in guitar tone modeling: EGTT replaces a recording's tone with that of a reference, while EGTR recovers the dry direct-input (DI) signal from a wet, processed recording. Despite their highly related nature, prior work has addressed them independently, and both works have yet to achieve satisfactory results. In this paper, we propose EG-VAE, a unified framework that jointly models EGTT and EGTR by disentangling frame-level content and global tone representations from wet recordings with a variational autoencoder. EGTT is achieved by recombining a source's content with a reference's tone, while EGTR is attained by a novel tone masking objective that enforces content-tone disentanglement during training and realizes the removal procedure at inference. To improve transfer to tones unseen in training, a second training stage shapes a smooth tone space through variational sampling and audio-effects augmentation. Experimental results from both objective and subjective evaluations demonstrate that EG-VAE outperforms task-specific baselines on transfer and removal. Demos are available at https://guitar-tone-demo.vercel.app/.