arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04975cs.SD

SwanWeave:单阶段多任务指令引导的3D空间音频编辑

One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing

发表机构浙江大学 · 字节跳动
查看机构详情
  • Zhejiang University(浙江大学)
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

Ke Lei, Chenyuhao Wen, Yu Zhang, Wenxiang Guo, Changhao Pan, Sashuai Zhou, Yongshi Li, Ruiqi Li, Ruofan Hu, Haorui Xu, Xiang Yin, Zhou Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

SwanWeave是首个单阶段多任务指令引导的3D FOA空间音频编辑框架,采用SE-MoE与SPO方法,在所有任务上优于现有通用音频编辑器和空间音频基线。

中文摘要 AI 辅助

空间音频编辑根据用户指令修改现有声场,同时保留场景其余部分,与传统音频编辑不同,它必须结合一阶Ambisonic(FOA)波形中的音频事件、空间信息、动态变化和环境信息进行联合推理。现有语言引导编辑器主要针对传统音频或依赖顺序操作,因此无法直接支持复杂3D空间指令的单阶段编辑。我们提出SwanWeave,这是首个用于指令引导的3D FOA空间音频编辑的单阶段多任务框架。我们使用可控房间模拟从开源语音和音效语料库构建配对FOA监督,覆盖四个编辑轴上的十多个单操作和复合任务。为处理这种异构编辑空间,SwanWeave采用双级路由的空间编辑混合专家(SE-MoE),为复合指令选择任务感知的专家组合,为局部编辑决策采用帧级路由/空专家。我们进一步引入空间偏好优化(SPO),一种基于直接偏好优化(DPO)的对齐目标,具有编辑特定的负样本目标,并采用分阶段训练以提升自然语言接地性。实验表明,SwanWeave在所有任务上均优于现有通用音频编辑器和空间音频基线。空间音频编辑演示可在此httpsURL查看,代码可在此httpsURL获取。

英文摘要

Spatial audio editing modifies an existing soundfield according to a user's instruction while preserving the rest of the scene. Unlike conventional audio editing, it must reason jointly about audio events, spatial information, dynamic changes, and environmental information in first-order Ambisonic (FOA) waveforms. Existing language-guided editors mainly target conventional audio or rely on sequential operations, and therefore do not directly support one-stage editing for complex 3D spatial instructions. We present SwanWeave, the first one-stage multi-task framework for instruction-guided 3D FOA spatial audio editing. We build paired FOA supervision from open-source speech and sound-effect corpora using controllable room simulation, covering more than ten single-operation and compound tasks across the four editing axes. To handle this heterogeneous edit space, SwanWeave uses Spatial Edit Mixture-of-Experts (SE-MoE) with dual-level routing, selecting task-aware expert combinations for compound instructions and frame-level routed/null experts for local edit decisions. We further introduce Spatial Preference Optimization (SPO), a Direct Preference Optimization (DPO)-based alignment objective with edit-specific negative targets, and adopt staged training to improve natural-language grounding. Experiments show that SwanWeave achieves better editing quality than existing general audio editors and spatial audio baselines across all tasks. Spatial audio editing demos can be found at https://swanaigc.github.io/#swanweave, code can be found at: https://github.com/MM-Speech/SwanWeave.

↑