arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04142cs.SDcs.AIcs.LG

InvFlowFD:基于流匹配逆变换的无参考且无背景集的感知音乐质量指标

InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

Alon Ziv, Harel Pogoda, Yossi Adi

AI总结:

本研究提出InvFlowFD,利用预训练流匹配骨干网络,通过流逆变换对比先验分布,实现无参考、无背景集的音乐感知质量评估,其与人类感知高度相关且更灵活。

AI中文摘要:

现有的无参考音乐感知质量评估方法无需成对的噪声-干净数据,但仍依赖于背景集来计算干净音频样本的聚合统计量。本研究提出一种消除该需求的新方法,仅使用预训练的流匹配(Flow Matching)骨干网络实现无背景集、无参考的质量估计。研究表明,通过简单欧拉积分实现的无条件流匹配逆变换足以检测各类人工失真,并能根据人类感知判断准确对音乐生成模型进行排名。我们引入InvFlowFD,该方法执行流逆变换并将一组逆变换样本与先验分布进行比较。我们通过定量分析和全面的人类研究将我们的方法与现有工作进行评估。结果表明,InvFlowFD与人类对声音失真的感知以及生成模型的质量高度相关,同时比现有指标更灵活、限制更少。

英文摘要:

Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples. In this work, we propose a novel approach that eliminates this requirement, achieving background-set-free and reference-free quality estimation using only a pre-trained Flow Matching backbone. We demonstrate that unconditional Flow Matching inversion via simple Euler integration is sufficient to detect various artificial distortions and accurately rank music generation models against human perceptual judgments. We introduce InvFlowFD, which performs flow inversion and compares a group of inverted samples to the prior distribution. We evaluate our method against prior work, quantitatively and with a thorough human study. Results suggest that InvFlowFD is highly correlated with human perception of sound distortions, as well as generative models' quality, while being more flexible and less restrictive than existing metrics.

补充信息

↑