arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TORUS:面向统一音频模型的渲染-理解自一致性测试

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji, Dinesh Manocha

arXiv 2607.28896首次发表:更新:

发表机构

Centific Global Solutions Inc.; University of Maryland(森蒂菲克全球解决方案公司; 马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出首个针对音频原生统一模型的自一致性测试TORUS,评估发现现有统一音频模型自一致性有限,在音频编辑任务上表现不佳,级联基线的自一致性表现优于最佳统一模型。

AI 中文摘要

能够进行音频理解、音频生成以及日益增多的音频编辑的统一音频模型正在迅速普及。然而,关于它们的一个基本问题仍未得到解答:统一模型的两个任务分支对同一段音频的判断是否一致?当前的做法是在专门的基准上分别评估每种能力,从未探究模型是否能理解自身的生成内容。我们提出了TORUS,这是首个针对音频原生统一模型的自一致性测试。TORUS包含48个三阶段自一致性测试,共432道六选一问题,覆盖五个任务系列中的语音、声音和音乐。我们全面评估了五个开源统一模型,以及结合最先进的专门生成、编辑和理解模型的级联基线。表现最佳的统一模型答对50.5%的问题,级联基线的正确率为63.2%,随机猜测的概率下限为16.7%。模型在音频编辑任务上表现吃力。在评估的音频模型(包括专门模型和统一模型)中,我们观察到有限的自一致性,因此将自一致性定位为未来音频系统的一项必要测试。

英文摘要

Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound and music across five task families. We holistically evaluate five open unified models alongside a Cascaded Baseline that combines state-of-the-art specialized generation, editing and understanding models. The best unified model answers 50.5% of questions against the Cascaded Baseline's 63.2% and a 16.7% chance floor. Models struggle on audio editing. Among the evaluated audio models (specialized and unified), we observe limited self-coherence, and thus position self-coherence as an essential test for future audio systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑