发表机构
Hangzhou Xiaoying Innovation Technology Co., Ltd. (Rythmix AI)(杭州小影创新科技有限公司(节奏混合人工智能))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人工智能生成翻唱歌曲的诊断评估,提出五维诊断框架,通过对30首翻唱歌曲分析发现和声进行和编排错误率高,各特征间关联有限,结果表明低级和符号摘要不能替代情境感知音乐判断。
AI 中文摘要
人工智能生成的翻唱歌曲常因局部音乐错误而失败,全局质量评分无法定位这些错误。我们提出了一个五维诊断框架,涵盖旋律音高、和声进行、调性一致性、风格一致性和编排/制作质量。基准包含由6个系统从5首源歌曲生成的30首翻唱歌曲,有专家严重程度评级和9个符号或声学特征。和声进行和编排的严重错误率最高,调性一致性保存较好。六个翻唱歌曲在调性一致性可接受的情况下存在严重和声错误。大跳比率与旋律评级有一定关联,但在多次测试后无特征相关性幸存。一个可解释的百分位数规则试点在16个维度级比较中也未能可靠地优于固定多数基线。结果将有用的诊断证据与可靠的自动评分区分开来:低级和符号摘要可揭示特定症状,但不能取代情境感知音乐判断。
英文摘要
AI-generated covers often fail through local musical errors that a global quality score cannot locate: the vocal contour may remain recognizable while the accompaniment uses the wrong harmonic function, or the output may stay in key while the arrangement remains incomplete. We present a five-dimensional diagnostic framework covering melodic pitch, harmonic progression, key consistency, style consistency, and arrangement/production quality. The benchmark contains 30 covers generated from 5 source songs by 6 systems, with expert severity ratings and 9 symbolic or acoustic features. Harmonic progression and arrangement had the highest severe-error rates (53% and 47%), whereas key consistency was better preserved. Six covers combined acceptable key consistency with severe harmonic errors. Large-leap ratio had a nominal association with melodic ratings (Spearman rho = -0.429, uncorrected p = 0.018), but no feature correlation survived the nine-test multiplicity reference. An interpretable percentile-rule pilot likewise failed to outperform a fixed majority baseline reliably across 16 dimension-level comparisons. The results separate useful diagnostic evidence from dependable automatic scoring: low-level and symbolic summaries can expose particular symptoms, but they do not replace context-aware musical judgment.
Comments16 pages, 3 figures, 8 tables. Code and analysis materials: https://github.com/TiaaL/songecho-cover-metrics