面向基于大语言模型的歌曲歌词性内容量化方法
Towards an LLM-based method for quantifying the sexual content in song lyrics
浏览论文内容
中文总结 AI 辅助
本文提出一种基于大语言模型的可复现方法,量化歌曲歌词主题内容(含性内容),并将其应用于12位雷鬼顿艺术家的1259首歌曲,还发布相关代码与语料库供他人复现或扩展。
中文摘要 AI 辅助
雷鬼顿(Reggaeton)是全球受众最广的音乐类型之一,其歌词通常被认为高度性化。这一说法主要基于定性研究和小规模定量研究。本文有两个目标:第一,提出一种可复现的方法,利用大语言模型(LLM)沿多个独立维度量化歌曲歌词的主题内容,该方法不仅限于性内容;第二,将其应用于包含12位雷鬼顿艺术家在2002年至2025年间发布的1259首歌曲的语料库。分析涵盖四个主题:数据集特征描述、艺术家间的比较、各维度随时间的变化分析,以及本文的性明确性评分与Spotify自身明确性标记的比较。我们发布了数据收集代码、评分提示和语料库,以便其他研究人员可复现该方法或将其应用于自身的歌词数据集。
英文摘要
Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mostly on qualitative studies and on small-scale quantitative ones. This paper has two goals. First, we present a reproducible method that uses a large language model to quantify thematic content in song lyrics along several independent dimensions. The method is not restricted to sexual content. Second, we apply it to a corpus of 1,259 songs by 12 reggaeton artists released between 2002 and 2025. The analysis covers four topics: a dataset characterization, a per-artist comparison, an analysis of how the dimensions change over time, and a comparison between our sexual-explicitness score and Spotify's own explicit flag. We release the data collection code, the scoring prompt, and the corpus, so that other researchers can replicate the approach or apply it to their own lyrics datasets.