发表机构
DePaul University; Hampton University(德保罗大学; 汉普顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Project Qualia通过分析大规模收听会话数据,训练Song2Vec模型并采用艺术家残差方法,成功从嵌入中恢复了独立于艺术家身份的体验性音乐结构,验证了该结构的可学习性。
AI 中文摘要
本报告展示了Project Qualia的结果,该项目是一项持续进行的工作,旨在确定歌曲之间的体验性相似性(一种未被流派或元数据分类法捕获的结构)是否可以从真实收听行为中恢复。我们构建了一个大规模的收听会话数据集,包含通过this http URL API从9,396名用户收集的12.9亿次收听记录,并通过预处理管道缩减为5.316亿次训练收听记录,涵盖2,860万个会话。在此语料库上,我们训练了一个skip-gram Word2Vec模型(Song2Vec),将会话视为句子,将每首曲目视为标记。正如预期,生成的嵌入空间主要由艺术家身份主导,这是会话中单一艺术家连续播放的结果。为了测试更微妙的、与艺术家无关的信号,我们开发了一种艺术家残差程序:从曲目的嵌入中减去每个艺术家的质心,并评估剩余部分是否保留结构。跨艺术家的平均余弦相似度从原始嵌入空间中的0.2487降至残差空间中的0.0005,然而在残差空间中,有4,577对跨艺术家曲目对保持了余弦相似度≥0.70,形成了连贯的基于流派和时代的聚类,包括trip-hop、1990年代grunge、2020年主流流行音乐,以及跨作曲家的古典钢琴曲对,其余弦相似度高达0.95。这些结果证实训练数据包含独立于艺术家身份的体验性结构,为设计直接学习该体验性层的架构奠定了实证基础。
英文摘要
This report presents results from Project Qualia, an ongoing effort to determine whether experiential similarity between songs, a structure not captured by genre or metadata taxonomies, can be recovered from real listening behavior. We constructed a large-scale dataset of listening sessions, comprising 1.29 billion scrobbles collected from 9,396 users via the Last.fm API and reduced through a preprocessing pipeline to 531.6 million training scrobbles across 28.6 million sessions. On this corpus, we trained a skip-gram Word2Vec model (Song2Vec), treating each session as a sentence and each track as a token. As anticipated, the resulting embedding space was dominated by artist identity, a consequence of single-artist runs within sessions. To test for a subtler, artist-independent signal, we developed an artist-residual procedure: subtracting each artist's centroid from its tracks' embeddings and evaluating whether the remainder retained structure. Mean cross-artist cosine similarity fell from 0.2487 in raw embedding space to 0.0005 in residual space, yet 4,577 cross-artist track pairs retained cosine similarity $\ge 0.70$ in residual space, forming coherent genre- and era-based clusters, including trip-hop, 1990s grunge, 2020 mainstream pop, and cross-composer classical piano pairs at cosine similarity up to 0.95. These results confirm that the training data contains experiential structure independent of artist identity, establishing an empirical basis for an architecture designed to learn this experiential layer directly.
Comments10 pages, 3 figures, 4 tables