发表机构
Dolby Laboratories(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PLACE通过条件嵌入和低秩变换扩展AudioX,实现多模态双耳音频生成,在FAIR-Play和BEWO-1M上提升空间一致性。
AI 中文摘要
我们提出了PLACE,一种将预训练的任意到音频模型AudioX扩展为从文本、视频和可选音频提示的任意组合进行双耳生成的方法。PLACE通过感知编码器核心特征增强视频条件,对齐文本和视频表示以推导空间线索,并对生成的潜在表示应用依赖于条件的低秩变换。该适配器通过解码音频的耳间电平和时间差目标进行监督。在MRSAudio上训练,PLACE在FAIR-Play上相比ViSAGe改善了大多数指标,并在BEWO-1M单静态测试分割上相比SpatialSonic取得了更高的SpatialCLAP分数。听众评估倾向于PLACE用于视频到音频和分布外文本到音频生成,展示了灵活的多模态控制和改进的空间一致性。
英文摘要
We present PLACE, a method that extends the pretrained any-to-audio model AudioX for binaural generation from arbitrary combinations of text, video, and optional audio prompts. PLACE augments video conditioning with Perception Encoder Core features, aligns text and video representations to derive spatial cues, and applies a conditioning-dependent low-rank transformation to the generated latent. The adapter is supervised via decoded-audio interaural level and time difference objectives. Trained on MRSAudio, PLACE improves most metrics over ViSAGe on FAIR-Play and achieves an improved SpatialCLAP score over SpatialSonic on the BEWO-1M Single Static test split. Listener evaluations favor PLACE for video-to-audio and out-of-distribution text-to-audio generation, demonstrating flexible multimodal control and improved spatial consistency.
CommentsSubmitted to ICASSP 2027