arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29987cs.SDeess.AS

生成式音乐模型在情绪条件跟随方面的表现如何?

How Well Do Generative Music Models Follow Emotion Conditioning?

Morteza Heydari

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建统一评估流程,以GTZAN数据集为基础,评估Stable Audio Open等三个生成式音乐模型的情绪跟随表现,发现文本条件生成更优,效价保留更可靠,强调直接评估情感可控性的重要性。

中文摘要 AI 辅助

近期的生成式音乐模型通过文本和音频条件提供了日益精细的控制,但它们遵循预期情绪线索的忠实程度仍是一个未解决的问题。我们针对生成音乐的情绪跟随问题,提出了一套统一的评估流程。我们使用GTZAN数据集中的全部1000首曲目,借助音频字幕模型DashengLM提取语义音频描述,同时利用音乐情感识别模型Music2Emotion估算源曲目的效价与唤醒度。我们将这些描述与排名靠前的情绪标签相结合,构建出感知情感的文本提示,并使用三个系统生成30秒的音频输出,分别是Stable Audio Open、MusicGen和InspireMusic,同时评估文本条件和音频条件下的生成效果。为衡量情绪跟随程度,我们计算生成音频的效价与唤醒度,并在效价-唤醒度空间中使用绝对误差和欧氏距离将其与源曲目进行比较。结果显示,文本条件下的生成效果始终优于音频条件,其中MusicGen(文本)和InspireMusic(文本)表现最佳,而音频条件的变体稳定性较差。我们还发现,效价的保留比唤醒度更可靠,且情绪跟随程度在不同音乐类型间存在显著差异。这些发现强调了直接评估情感可控性的重要性,而非仅依赖通用质量或提示相关性指标。

英文摘要

Recent generative music models offer increasingly fine-grained control through text and audio conditioning, yet how faithfully they follow intended emotional cues remains an open question. We address this gap with a unified evaluation pipeline for emotion-following in generated music. Using all 1000 tracks in GTZAN, we extract semantic audio descriptions with DashengLM, an audio captioning model, and estimate source-track valence and arousal with Music2Emotion, a music emotion recognition model. We construct affect-aware text prompts by combining descriptions with top-ranked emotion tags and generate 30-second outputs with three systems, Stable Audio Open, MusicGen, and InspireMusic, evaluating both text- and audio-conditioned generation. To measure emotion-following, we compute valence and arousal on generated audio and compare them with the source tracks using absolute error and Euclidean distance in valence-arousal space. Text-conditioned generation consistently outperforms audio conditioning, with MusicGen (text) and InspireMusic (text) achieving the best performance, while audio-conditioned variants prove less stable. We further find that valence is preserved more reliably than arousal and that emotion-following varies substantially across genres. These findings underscore the importance of evaluating affective controllability directly rather than relying solely on general quality or prompt-relevance metrics.

补充信息

↑