arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BackgroundMellow:一种用于叙事驱动的丰富电影音效生成的多模态凝聚框架

BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation

Ajitesh Jamulkar, Aritra Hazra

arXiv 2607.11364首次发表:更新:

发表机构

Indian Institute of Technology Kharagpur(印度理工学院卡拉格布尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

多模态AI中为文本叙事生成电影音效困难,BackgroundMellow框架将其视为编排与信号处理问题,通过主 - 专家架构、Tango2模型、背景音乐检索器及NLP模块实现,经实验验证在时间同步等方面有效。

AI 中文摘要

在多模态人工智能中,为长篇文本叙事生成沉浸式、同步且具有电影感的音频仍然是一项重大挑战。当前的文本到音频(TTA)框架虽能成功合成孤立音效,但在叙事连贯性、时间对齐和电影情感深度方面存在困难。我们提出了BackgroundMellow框架,将故事到音频的生成视为精确编排和信号处理问题。该框架通过主 - 专家代理架构在无真实标签情况下运行,将文本分解为精确且多层的音频线索,用合适的专家模型生成各类声音并叠加音效以创建统一且对齐的音频片段。我们的流程基于Tango2潜在扩散模型进行环境合成,并结合从专业配乐中挖掘的新型电影背景音乐检索器。为实现声音混合过程自动化,我们使用基于自然语言处理的模块根据叙事时间线预测精确音频参数。我们还通过对精心策划的YouTube电影预告片数据集进行最近邻检索,实证评估并展示了所提框架在时间同步、覆盖范围和频谱丰富度方面的有效性。

英文摘要

Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI. While current Text-to-Audio (TTA) frameworks successfully synthesize isolated sound effects, they struggle with narrative cohesion, temporal alignment, and cinematic emotional depth. We present BackgroundMellow, a framework that treats story-to-audio generation as a precise orchestration and signal processing problem. This framework is enabled without ground-truth through a master-specialist agent architecture that decomposes text into precise and multi-layered audio cues, generates each category of sounds with suitable specialist model, and superimposes the soundscapes to create a unified and aligned audio segment. Our pipeline is built over Tango2 latent diffusion model for environmental synthesis alongside a novel Cinematic BGM Retriever mined from professional soundtracks. To automate the sound mixing process, we use an NLP based module that predicts precise audio parameters, like start time, duration, and relative loudness, based on the narrative timeline. We further empirically evaluate and show the efficacy of the proposed framework leveraging nearest-neighbor retrieval against a curated dataset of YouTube cinematic trailers to measure temporal synchronization, coverage, and spectral richness.

Comments7 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑