从提示到描述:针对AI生成音乐语言的跨文化研究
From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music
浏览论文内容
中文总结 AI 辅助
该研究对比了TTM系统提示与音乐描述的语言差异,构建了音乐提示词汇分类法,发现提示与感知存在结构不对称,且不同文化背景的描述特征存在差异,引发对现有TTM系统文化适应性的思考。
中文摘要 AI 辅助
文本到音乐(TTM)生成系统允许用户通过自然语言提示创作音乐,但目前尚不清楚用于提示的描述性语言是否与用于总结或描述所听音乐的描述性语言一致。我们将200个真实世界的Udio提示与其生成的音频配对,并收集英语(n=70)和韩语(n=78)听众的自由形式描述,基于真实用户数据贡献了一种人类推导的音乐提示词汇分类法。利用该框架以及词级和向量级分析,我们发现了一种一致的结构不对称性:提示以类型(Genre)和故事/叙事语言为主。类型术语从提示到感知的传播最可靠,而叙事密集的提示是语义不一致的最强预测因素。初步跨文化比较进一步表明,不同听众群体在叙事、功能和情感维度上的描述特征存在差异,这引发了关于当前以英语为中心的聚合语料库训练的TTM系统是否能适应人们自然表达音乐想法的全部多样性的疑问。
英文摘要
Text-to-music (TTM) generation systems allow users to create music through natural language prompts, yet it is unclear whether the descriptive language used to prompt aligns with descriptive language used to summarize or describe heard music. We pair 200 real-world Udio prompts with their generated audio and free-form descriptions collected from English- (n = 70) and Korean-speaking (n = 78) listeners, and contribute a human-derived taxonomy of musical prompting vocabulary grounded in real user data. Using this framework, alongside word- and vector-level analyses, we find a consistent structural asymmetry: prompts are dominated by Genre and Story/Narrative language. Genre terms propagate most reliably from prompt to perception, while narrative-heavy prompts are the strongest predictor of semantic misalignment. A preliminary cross-cultural comparison further suggests that description profiles vary across listener populations along narrative, functional, and affective dimensions, raising questions about whether current TTM systems, trained on aggregated English-centric corpora, can accommodate the full diversity of how people naturally express musical ideas.