arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39552cs.SDcs.LG

编码器中的幽灵:歌词到歌曲生成中可解码的艺术家身份表示

Ghost in the Encoder: Decodable Artist Identity Representations in Lyrics-to-Song Generation

Arhan Vohra, Choenden Kyirong, Laura Ibáñez-Martínez, Martín Rocamora

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过探测ACE-Step 1.5模型内部激活,证明歌词可线性解码出艺术家身份表示,揭示歌词构成艺术家级条件通道,并展示潜在空间分析用于审计生成音乐模型隐式学习内容的方法。

中文摘要 AI 辅助

文本到歌曲生成模型可以被提示模仿特定艺术家或从其训练数据中复述完整歌曲。尽管这些现象已在小型数据集上通过行为方式得到记录,但关于可能引发这些现象的内部表示却知之甚少。先前关于生成音频的可解释性工作侧重于在模型激活中定位语义概念,如流派或拍号。在本工作中,我们表明,经过训练的模型可以仅从歌词中被探测出艺术家身份的线性可解码表示,而无需任何额外的标识符。通过对ACE-Step 1.5的受控案例研究,涵盖100位艺术家的2,000首歌曲,我们证明了与给定歌词集相关的艺术家可以在模型内部激活中被识别,并且这种条件信号在推理期间从歌词编码器传播到扩散主干。这些发现表明,歌词构成了一个艺术家级别的条件通道,而提示侧复制保障措施并未解决该通道。更广泛地说,我们的工作凸显了潜在空间分析如何用于审计生成音乐模型从其训练数据中隐式学习到的内容。

英文摘要

Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these phenomena have been documented behaviorally on small datasets, little is known about the internal representations that may give rise to them. Prior interpretability work on generative audio has focused on locating semantic concepts such as genre or time signature within model activations. In this work, we show that a trained model can be probed for linearly decodable representations of artist identity from song lyrics alone, without any additional identifiers. Through a controlled case study of ACE-Step 1.5 spanning 2,000 songs across 100 artists, we demonstrate that the artist associated with a given set of lyrics can be identified within the model's internal activations, and that this conditioning signal propagates from the lyric encoder to the diffusion backbone during inference. These findings indicate that lyrics constitute an artist-level conditioning channel not addressed by prompt-side replication safeguards. More broadly, our work highlights how latent-space analysis can be used to audit what generative music models have implicitly learned from their training data.

补充信息

↑