arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从分类到推荐:音频嵌入模型在基于内容的音乐推荐中应用的实证分析

From Classification to Recommendation: Empirical Analysis of Audio Embedding Models Application for Content-Based Music Recommendation

Qingrui Li, Haowei Lou, Chengkai Huang, Quan Z. Sheng, Lina Yao

arXiv 2608.06928首次发表:更新:

AI 中文总结

该研究评估六种音频编码器在三类音乐推荐系统中的效果,分析残差量化设计的影响,为现代音乐推荐系统的音频编码器选择及语义ID设计提供实用指导。

AI 中文摘要

从大规模语料库学习得到的预训练音频表示模型在音频分类与理解任务中已取得优异性能。然而,现有大多数模型针对掩码预测、对比学习或音频-文本对齐等目标优化,其生成的表示空间未必适配推荐系统。与分类任务不同,音乐推荐系统需捕捉由主观及行为依赖的听众偏好所塑造的项目关系。尽管预训练音频嵌入已在传统推荐系统中得到探索,但它们在快速兴起的生成式推荐系统范式中的有效性仍未得到充分研究。为填补这一空白,我们针对三类音乐推荐系统(基于内容的、序列式的以及基于语义ID的生成式推荐系统),系统评估了六种代表性音频编码器。我们进一步研究了残差量化设计(包括码本宽度、量化深度及保留的语义ID前缀)对推荐相关信息保留的影响。在两个音乐推荐数据集上的实验表明,当直接使用预训练嵌入几何时,音频-文本对齐及音乐领域表示通常更有效;而基于交互的序列式训练则会显著降低编码器之间的性能差异。我们还发现,增加语义ID容量并不总能提升生成式推荐系统的性能,反而可能引发显著的不稳定性。这些发现为现代音乐推荐系统中音频编码器的选择及音频衍生语义ID的设计提供了实用指导。

英文摘要

Pretrained audio representation models learned from large-scale corpora have achieved strong performance in audio classification and understanding. However, most existing models are optimized for objectives such as masked prediction, contrastive learning, or audio-text alignment, which do not necessarily produce representation spaces well-suited to recommender systems. Unlike classification, music recommender systems must capture item relationships shaped by subjective and behavior-dependent listener preferences. Although pretrained audio embeddings have been explored in conventional recommender systems, their effectiveness in the rapidly emerging paradigm of generative recommender systems remains underexplored. To address this gap, we systematically evaluate six representative audio encoders across three types of music recommender systems: content-based, sequential, and Semantic-ID-based generative recommender systems. We further investigate how residual-quantization design, including codebook width, quantization depth, and retained Semantic-ID prefixes, affects the preservation of recommendation-relevant information. Experiments on two music recommendation datasets show that audio-text-aligned and music-domain representations are generally more effective when pretrained embedding geometry is used directly, whereas interaction-based sequential training substantially reduces performance differences among encoders. We also find that increasing Semantic-ID capacity does not consistently improve generative recommender systems and may introduce substantial instability. These findings provide practical guidance for selecting audio encoders and designing audio-derived Semantic IDs for modern music recommender systems.

Comments9 pages, 4 tables, 2 figures. v2: Added corresponding-author and contact information, consolidated a duplicate reference, completed bibliographic metadata, and corrected the post-reference figure layout. Technical content and results are unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑