AI 中文总结
CloudEar是整合三类证据的音乐评估成对框架,含三类专家组件,在音乐评估的成对偏好准确率与分数相关性上优于现有基线。
AI 中文摘要
评估生成音乐需建模感知质量、技术音频质量及提示对齐。现有方法常忽略信号层面缺陷,生成压缩的质量分数,且遗漏文本提示中的细粒度音乐属性。我们提出CloudEar,一种用于评估音乐性与提示对齐的成对框架,包含三个组件:i)用于歌曲质量评估的感知专家;ii)用于学习及信号层面质量分析的诊断专家;iii)基于音乐属性实现提示-歌曲对齐的跨模态专家。实验表明,CloudEar在成对偏好准确率与分数相关性上均优于对比基线。
英文摘要
Evaluating generated music requires modeling perceptual quality, technical audio quality, and prompt alignment. Existing methods often overlook signal-level defects, produce compressed quality scores, and miss fine-grained musical attributes in text prompts. We propose CloudEar, a pairwise framework for evaluating musicality and prompt alignment. It comprises three components: i) a Perceptual Expert for song-quality assessment; ii) a Diagnostic Expert for learned and signal-level quality analysis; and iii) a Cross-modal Expert for prompt-song alignment based on musical attributes. Experiments show that CloudEar outperforms the compared baselines in pairwise preference accuracy and score correlation.
CommentsAccepted at the Late-Breaking/Demo session of ISMIR 2026