arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个分数够吗?用时间分数曲线评估歌曲的演唱质量

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

Yishan Lv, Jing Luo, Xinyu Yang, Zhizheng Wu

arXiv 2607.16599首次发表:更新:

发表机构

Xi’an Jiaotong University; The Chinese University of Hong Kong, Shenzhen(西安交通大学; 香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对全长歌曲演唱质量评估难题,提出SongSQA两阶段框架。第一阶段用伪标签训练段分数预测器,第二阶段聚合器整合特征与分数生成嵌入并捕捉联系,有效提升评估效果,在KTAU上相对提高13.95%。

AI 中文摘要

演唱质量评估(SQA)在实际多媒体应用和音乐人工智能系统中变得越来越重要,但现有研究主要集中在短演唱片段,对全长歌曲的评估还不够。全长歌曲SQA需要对演唱质量在不同音频段的变化以及这些局部变化如何影响整体演唱表现评估进行建模。此外,段级注释的稀缺使得有效监督具有挑战性。为应对这些挑战,我们提出了SongSQA,这是一个用于全长歌曲SQA的两阶段框架。第一阶段,使用预训练教师模型生成的伪标签训练段分数预测器,无需手动段注释即可进行段级演唱质量预测。第二阶段,歌曲质量聚合器将段特征和预测的段分数集成到统一的段嵌入中,并使用可学习的歌曲嵌入和自注意力来捕捉段级演唱表现与整体歌曲质量之间的联系。实验结果证明了SongSQA对全长歌曲SQA的有效性,在KTAU上比最强基线相对提高了13.95%,同时在所有数据集上持续改进其他评估指标。

英文摘要

Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length song SQA requires modeling how singing quality varies across different audio segments and how these local variations influence the overall evaluation of vocal performance. Moreover, the scarcity of segment-level annotations makes effective supervision challenging, as directly assigning a single overall score label to every segment tends to treat different segment qualities as equivalent. To address these challenges, we propose SongSQA, a two-stage framework for full-length song SQA. In the first stage, a Segment Score Predictor is trained with pseudo labels generated by a pre-trained teacher model, enabling segment-level singing quality prediction without requiring manual segment annotations. In the second stage, a Song Quality Aggregator integrates segment features and predicted segment scores into unified segment embeddings, and employs a learnable song embedding together with self-attention to capture the connection between segment-level vocal performance and overall song quality. In this way, SongSQA dynamically aggregates critical quality cues across the song to produce a holistic quality prediction, while also generating a temporal segment-level quality curve. Experimental results demonstrate the effectiveness of SongSQA for full-length song SQA, achieving up to a 13.95% relative improvement in KTAU over the strongest baseline, while consistently improving other evaluation metrics across all datasets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑