arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AudioScape-TTA:用于细粒度文本到音频评估的结构化声景基准

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

Jinting Wang, Yuguang Yang, Shengyu Li, Yan Rong, Shan Yang, Xiaoda Yang, Li Liu

arXiv 2608.04479首次发表:更新:

AI 中文总结

针对现有TTA评估基准缺乏细粒度语义洞察的问题,提出AudioScape-TTA基准及基于评分标准的评估框架,通过2258个音频-文本对等数据验证TTA模型的多项局限,其评估更贴合人类判断。

AI 中文摘要

文本到音频(TTA)生成近期在根据自然语言描述合成真实音频方面取得了显著进展,但判断生成音频是否忠实满足复杂文本指令仍具挑战性。现有基准主要依赖全局相似度指标,对细粒度语义故障的洞察有限。为解决这一局限,我们引入AudioScape-TTA,这是一个结构化且感知复杂度的细粒度TTA评估基准。AudioScape-TTA通过模态感知语义结构表征真实声景,并利用事件密度和结构复杂度刻画生成复杂度。基于这些标注,我们提出了基于评分标准的音频接地评估框架,通过细粒度语义标准验证事件实现、声学属性和语音内容。该基准包含2258个音频-文本对及25707个二元问答评分标准,支持对TTA系统的可扩展且可解释的分析。对13个代表性开源TTA模型的实验显示,其在细粒度属性控制、语音内容保留及组合声景生成方面存在持续局限;人工验证进一步表明,我们的基于评分标准的评估比传统全局相似度指标更贴合人类语义判断。

英文摘要

Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated audio faithfully satisfies complex textual instructions remains challenging. Existing benchmarks mainly rely on global similarity metrics, providing limited insight into fine-grained semantic failures. To address this limitation, we introduce \textbf{AudioScape-TTA}, a structured and complexity-aware benchmark for fine-grained TTA evaluation. AudioScape-TTA represents realistic soundscapes through modality-aware semantic structures and characterizes generation complexity using event density and structural complexity. Based on these annotations, we propose a rubric-based audio-grounded evaluation framework that verifies event realization, acoustic attributes, and speech content through fine-grained semantic criteria. The benchmark contains 2,258 audio-text pairs with 25,707 binary QA rubrics, enabling scalable and interpretable analysis of TTA systems. Experiments on 13 representative open-source TTA models reveal persistent limitations in fine-grained attribute control, speech-content preservation, and compositional soundscape generation. Human validation further demonstrates that our rubric-based evaluation achieves stronger alignment with human semantic judgments than conventional global similarity metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑