arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03742cs.SDcs.AI

基于AI的音效生成:跨输入模态生成模型的叙事性综述

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins, Adam Dubrowski

AI总结:

本章综述了跨文本、视觉等输入模态的AI音效生成模型,考察30篇文献,指出模型性能提升的同时存在时间同步局限等挑战,推动音效生成向自适应系统发展。

AI中文摘要:

音效在数字应用中对传递动作、事件和环境提示至关重要,通常需要高度的可变性和情境适应性。人工智能(AI)驱动的音频生成模型正迅速普及,有望改变各类应用中声音的合成与使用方式。针对这一发展态势,本章综述并分析了近期用于音效合成的AI生成模型,重点关注不同输入模态(文本、视觉、音频及多模态)如何影响生成音频的质量、可控性和情境相关性。研究考察了来自Google Scholar、IEEE Xplore和ACM Digital Library的30篇同行评审文章,探讨了过去五年AI生成模型的演进。结果显示,多个模型实现了最先进的性能,在各类任务中生成了高保真、语义对齐且时间连贯性日益提升的音效。然而,尽管取得了这些进展,本综述仍发现了持续存在的挑战,包括复杂多事件场景下的时间同步局限、客观指标与人类感知之间的差距,以及可控性与生成多样性之间的权衡。总体而言,本章强调AI驱动的音效生成正朝着更具适应性、可扩展性和情境感知的系统发展,对未来的声音设计工作流和交互式媒体应用具有重要意义。

英文摘要:

Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform the way sound is synthesized and used across various applications. In response to this growing momentum, this chapter reviews and analyzes recent AI-based generative models for sound effect synthesis, with a focus on how different input modalities (text, visual, audio, and multimodal) affect the quality, controllability, and contextual relevance of the generated audio. It examines 30 peer-reviewed articles sourced from Google Scholar, IEEE Xplore, and the ACM Digital Library, exploring the evolution of AI generative models over the past five years. The results show that multiple models achieved state-of-the-art performance, producing high-fidelity, semantically aligned, and increasingly temporally coherent sound effects across tasks. However, despite these advances, the review identifies persistent challenges, including limitations in temporal synchronization for complex multi-event scenarios, gaps between objective metrics and human perception, and trade-offs between controllability and generative diversity. Overall, the chapter highlights that AI-driven sound effect generation is progressing toward more adaptive, scalable, and context-aware systems, offering significant implications for future sound design workflows and interactive media applications.

补充信息

↑