arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SoundscapeAgent:用于可控合成和可扩展音频-语言监督的智能音景构建

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

Hao Zhang, Yiwen Zhao, Yixuan Zhang, Yiwen Shao, Steve Yves

arXiv 2607.21857首次发表:更新:

发表机构

Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出智能音景构建框架SoundscapeAgent,基于大语言模型的智能体将用户意图转化为场景计划,经多步骤实现可控音频合成与可扩展音频-语言数据构建,在生成性能及下游音频推理方面表现出色。

AI 中文摘要

我们提出了一个用于可控合成音频生成的智能音景构建框架,该框架明确了通常由单步文本到音频模型隐式处理的场景规划、源选择、时间布局和渲染步骤。基于大语言模型的智能体将用户意图转换为可执行的场景计划,通过检索和按需生成获取资产,渲染可控的多事件混合,并导出对齐的场景元数据。该框架还通过用户引导的工具选择和可编辑的场景计划支持人在回路交互。这些组件共同提供了一种可检查且可重复使用方法,用于可控音景合成和可扩展音频-语言数据构建。听众研究和客观指标表明,与文本到音频基线相比,其生成性能具有竞争力,而使用智能体生成的数据训练的模型在下游音频推理中始终优于仅使用真实数据的基线。代码、演示和听力测试结果可在该https网址获取

英文摘要

We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, and rendering steps typically handled implicitly by single-shot text-to-audio models. An LLM-based agent converts user intent into an executable scene plan, acquires assets through retrieval and on-demand generation, renders controllable multi-event mixtures, and exports aligned scene metadata. The framework also supports human-in-the-loop interaction through user-guided tool selection and editable scene plans. Together, these components provide an inspectable and reusable approach to controllable soundscape synthesis and scalable audio-language data construction. Listener studies and objective metrics demonstrate competitive generation performance against text-to-audio baselines, while models trained with agent-generated data consistently outperform real-only baselines in downstream audio reasoning. Code, demos, and listening-test results are available at https://haozhang6720.github.io/SoundscapeAgentDemoPage/.

CommentsSubmitted to SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑