arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06106cs.CY

声音的算法扁平化:AI音乐同质化的计算证据与正义意涵

The Algorithmic Flattening of Sound: Computational Evidence and Justice Implications of AI Music Homogenization

Zoe Slendebroek, Danaé Metaxa

首次发表
浏览论文内容

中文总结 AI 辅助

本文审计Suno和Lyria 3在四类音乐中的同质化趋势,发现二者存在不同的同质化模式,AI与人类曲目可被MIR特征近乎完美区分,该现象关乎音乐风格的可识别性与经济回报等正义问题。

中文摘要 AI 辅助

本文研究大规模生成式音乐系统是否相较于人类创作的音乐存在可测量的音乐同质化现象,并从正义视角阐释该现象的重要性。我们对两个商业部署的系统Suno和Lyria 3在Afrobeats、K-pop、Dance Pop和Heavy Metal四个音乐流派中开展审计,针对每个系统和流派各生成100首曲目,采用72个音乐信息检索(MIR)特征以及离散度、冗余度和可分性的多项诊断指标,将同质化定义为标准计算音频特征(包括节奏与 timing、音色/频谱形态、动态范围)在流派内部及流派边界间的声学变异减少。我们还仅以流派名称为提示词生成曲目,无额外指令,以揭示每个系统的默认音乐倾向。结果显示两种结构不同的同质化趋势:Lyria降低了流派内部的声学多样性,而Suno则消除了流派间的声学差异,却未压缩流派内部的分布范围。两个系统均未忠实地遵循用户提示词,表明观察到的模式反映的是学习到的先验而非提示词约束。两个系统并未收敛于共同的声学特征,彼此间的声学距离比两组随机人类子样本通常的距离更大。尽管如此,标准分类器仅借助MIR特征就能近乎完美地区分AI与人类创作的曲目。我们认为这些模式并非审美层面的猎奇,而是与正义相关的状况,随着生成式输出大规模传播,它们会塑造哪些音乐风格变得可识别、有价值并获得经济回报。

英文摘要

This paper audits whether large-scale generative music systems exhibit measurable musical homogenization relative to human-produced music, and develops a justice-centered account of why this matters. We audit two commercially deployed systems (Suno and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, and Heavy Metal). For each system and genre, we generate 100 tracks and compare them against human corpora of equal size, using 72 music information retrieval (MIR) features and multiple diagnostics of dispersion, redundancy, and separability. We define homogenization as reduced acoustic variation in standard computational audio features including rhythm and timing, timbre/spectral shape, and dynamics, both within genres and across genre boundaries. We also generate tracks using only a genre name as the prompt, with no additional instructions, to reveal each system's default musical tendencies. The results show two structurally distinct homogenizing tendencies. Lyria reduces within-genre acoustic diversity, while Suno collapses the acoustic distinctions between genres without compressing within-genre spread. Neither system follows user prompts faithfully, indicating that the observed patterns reflect learned priors rather than prompt constraints. The two systems do not converge on a common acoustic profile and are more acoustically distant from each other than two random human subsamples would typically be. Nevertheless, a standard classifier distinguishes AI from human tracks near-perfectly on MIR features alone. We argue that these patterns matter not as an aesthetic curiosity but as a justice-relevant condition, shaping which musical styles become legible, valued, and economically rewarded as generated outputs increasingly circulate at scale.

补充信息

↑