arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysWave:用于可控空间音频生成的物理引导潜在扩散模型

PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation

Lingfeng Yao, Chenpei Huang, Xingke Yang, Ziye Geng, Changqing Luo, Hao Wang, Jiang Liu, Miao Pan

arXiv 2608.29549首次发表:更新:

发表机构

University of Houston; Stevens Institute of Technology; Waseda University(休斯顿大学; 史蒂文斯理工学院; 早稻田大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhysWave是一种物理引导潜在扩散模型,通过统一控制与加入声学先验,生成空间一致的文本到FOA音频,还构建了30万段FOA数据集,可用于推理时的空间优化。

AI 中文摘要

文本到空间音频生成(如文本到一阶Ambisonics(FOA))为价值数十亿美元的游戏和电影行业创建空间音频提供了便捷方式。然而,现有的文本到FOA方法大多是数据驱动的,可能产生违反声源方向与距离间声学关系的音频,且它们将描述性控制与参数化控制分离,迫使用户在易用性与精确性间权衡。本文提出PhysWave,一种用于可控文本到FOA生成的物理引导潜在扩散模型。PhysWave通过共享航路点-字幕表示统一自然语言与轨迹控制,并在扩散训练中加入两种可微声学先验:球谐方向一致性与平方反比距离一致性。为支持动态空间生成,我们还构建了包含多样声音类别与声源轨迹的30万段音频的FOA数据集。大量结果表明,所提先验帮助PhysWave生成空间一致的FOA音频,同时保持有竞争力的音频质量。进一步分析显示,这些物理先验在训练期间提升空间一致性,还可作为推理时的引导用于无需训练的空间优化。

英文摘要

Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acoustic relations between source direction and distance. They also separate descriptive and parametric control, forcing users to trade usability for precision. In this paper, we present PhysWave, a physics-guided latent diffusion model for controllable text-to-FOA generation. PhysWave unifies natural-language and trajectory control through a shared waypoint-caption representation, and augments diffusion training with two differentiable acoustic priors: spherical-harmonic direction consistency and inverse-square distance consistency. To support dynamic spatial generation, we further construct a 300K-clip FOA dataset with diverse sound categories and source trajectories. Extensive results show that the proposed priors help PhysWave generate spatially consistent FOA audio while maintaining competitive audio quality. Further analyses show that these physics priors improve spatial consistency during training and can also be used as inference-time guidance for training-free spatial refinement.

CommentsAccepted by EMNLP 2026 Main Conference. Project website: https://lingfengyao.github.io/PhysWave/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑