arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09973cs.SDcs.AIcs.SYeess.ASeess.SPeess.SY

一种面向生产的声效生成评估框架

A Production-Oriented Framework for Evaluation of SFX Generation

Mélodie Desbos, Yara Bahram, Eric Granger, Mohammadhadi Shateri

首次发表
浏览论文内容

中文总结 AI 辅助

针对工业声音设计中声效生成评估问题,提出面向生产的评估框架,通过确定生产要求、两阶段协议及结合客观指标与人类研究,揭示不同基线的优势权衡,为声效变化评估建立协议并为设计音频生成管道提供基础。

中文摘要 AI 辅助

工业声音设计需要音频生成系统不仅能产生逼真音频,还能保留参考音频的感知特性、支持可控变化且在实际工作流程中保持高效。现有评估通常与文本到音频、无条件或特定任务设置相关,限制了对参考引导声效变化的评估。为此,提出一个面向生产的评估框架,用于异构音频生成和编辑方法的结构化比较。该框架确定了九个生产要求并明确考虑模型能力差异,通过两阶段协议进行评估,结合客观指标和人类感知研究。研究揭示了不同基线在不同生产需求下的优势和权衡。该框架为参考引导的声效变化建立了结构化评估和决策协议,为设计未来统一的工业音频生成管道提供了实践基础。

英文摘要

Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluations are usually tied to text-to-audio (TTA), unconditional, or task-specific settings, limiting assessment for reference-guided sound effects (SFX) variation. To address this gap, we present a production-oriented evaluation framework for structured comparison of heterogeneous audio generation and editing methods. Our framework identifies nine production requirements and explicitly accounts for differences in model capabilities, enabling comparison under a common production objective. A two-stage protocol is introduced: (1) a reference-guided audio-to-audio (ATA) variation task, in which all methods are evaluated under the same ESC-50 SFX adaptation setup, and (2) capability-specific analyses of native operations such as SFX morphing, temporal and energy alignment, inpainting, and targeted editing. This framework combines objective metrics (including FAD, ImageBind-based reference alignment, and diversity across generated variants), together with a human study of perceptual identity preservation and transient diagnosis. Our study reveals complementary strengths and trade-offs across baselines for different production needs. Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity while still supporting SFX morphing. Other baselines remain most suitable for specific editing operations. Our framework establishes a structured evaluation and decision protocol for reference-guided SFX variation and provides a practical basis for designing future unified industrial audio generation pipelines. Audio demos are on the accompanying web page.

补充信息

↑