arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10774cs.SD

AutoSynth:从音频和文本生成可编辑合成器程序的学习方法

AutoSynth: Learning to Generate Editable Synthesizer Programs from Audio and Text

Tristan Wu, Daniel Chin, Liwei Lin, Junan Zhang, Gus Xia

首次发表
浏览论文内容

中文总结 AI 辅助

AutoSynth是一种从音频或文本生成可编辑合成器程序的模型,经两阶段训练后,在合成器逆问题和文本驱动生成任务中表现具竞争力,无需配对标注或可微分合成器。

中文摘要 AI 辅助

音频生成模型可将自然语言描述转换为声音,但其输出通常是波形,音频质量受音频压缩限制,且输出难以直接编辑音符、音色参数或调制关系。我们提出AutoSynth,它将MIDI演奏事件、固定合成器参数和可变长度调制路径表示为合成器的统一序列,并通过音频条件自回归模型学习它们的依赖关系,单个模型支持两项任务:给定参考音频,模型直接预测合成器程序;给定文本,它使用预训练音频生成模型并将生成的音频转换为程序。训练分为两个阶段:第一阶段是对从少量原生预设自动构建的大规模音频-程序对进行监督学习,第二阶段是采用混合奖励的组相对策略优化,混合奖励结合语义相似度、音高相关特征、声学相似度和声音实用性。该流程既不需要配对的文本-目标-程序标注,也不需要可微分合成器。实验表明,AutoSynth可生成完整、可编辑的合成器程序,在合成器逆问题和文本驱动生成任务中均取得有竞争力的结果。音频演示和源代码可在指定URL获取。

英文摘要

Audio generation models can translate natural-language descriptions into sound, but their outputs are typically waveforms. Their audio quality is constrained by audio compression, and their outputs do not readily support direct edits to notes, timbral parameters, or modulation relationships. We present AutoSynth, which represents MIDI performance events, fixed synthesizer parameters, and variable-length modulation routes as a unified sequence for a synthesizer, and learns their dependencies with an audio-conditioned autoregressive model. A single model supports both tasks. Given reference audio, the model directly predicts a synthesizer program; given text, it uses a pretrained audio generation model and converts the generated audio into a program. Training consists of two stages: supervised learning on large-scale audio-program pairs automatically constructed from a small set of native presets, followed by group-relative policy optimization with a mixed reward combining semantic similarity, pitch-related features, acoustic similarity, and sound usefulness. The pipeline requires neither paired text-target-program annotations nor a differentiable synthesizer. Experiments show that AutoSynth produces complete, editable synthesizer programs and achieves competitive results in both synthesizer inversion and text-driven generation. Audio demos and source code are available at https://auto-synth.github.io/.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • New York University Shanghai(上海纽约大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑