arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STEMMA:面向大型音频语言模型的歌曲到音轨多音频推理

STEMMA: Song-to-Stem Multi-Audio Reasoning for Large Audio Language Models

Hoyeol Sohn, Wonil Kim, Keunhyoung Kim, Sangeun Kum, Taehyoung Kim, Dongjoo Moon, Theerasak Charoenchob, Teeratep Weerapang, Jongpil Lee, Juhan Nam

arXiv 2610.11884首次发表:更新:

发表机构

Neutune(Neutune)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出STEMMA多音频音乐问答框架,构建STEMMA-Bench与STEMMA-Instruct数据集,微调大型音频语言模型可提升多音频推理能力,同时保留单音频音乐理解能力。

AI 中文摘要

音乐理解通常需要比较片段并推理歌曲、段落与音轨(stem)之间的关系。然而,现有的大型音频语言模型(LALMs)和音乐问答数据集通常仅处理单条录音,或比较无已知制作关联的独立采样音轨。我们提出STEMMA,这是一个围绕制作来源构建的多音频音乐问答框架,制作来源指片段是否源自同一条音轨或段落,以及哪些音轨属于哪些混音。由于此类关系在传统的音频优先采样下较为稀疏,STEMMA采用关系优先的构建策略:它先指定目标关系,再从目录中查询满足该关系的片段和不满足的难负样本,标签直接由目录来源确定,而非由语言模型从元数据生成。我们构建了用于评估的STEMMA-Bench和音轨不相交的训练集STEMMA-Instruct。在STEMMA-Instruct上微调两个大型音频语言模型(LALMs)可提升多音频推理能力,在由目录直接确定的结构关系上提升最为显著,同时保留单音频音乐理解能力。

英文摘要

Music understanding often requires comparing excerpts and reasoning about relationships among songs, sections, and stems. However, existing large audio-language models (LALMs) and music question-answering datasets typically operate on single recordings or compare independently sampled tracks with no known production relationship. We introduce STEMMA, a multi-audio music question-answering framework built around production provenance: whether excerpts originate from the same track or section, and which stems belong to which mixtures. Because such relations are sparse under conventional audio-first sampling, STEMMA adopts a relation-first construction strategy: it first specifies a target relation and then queries the catalog for excerpts that satisfy it and hard negatives that do not. Labels are determined directly from catalog provenance rather than generated by a language model from metadata. We build STEMMA-Bench for evaluation and a track-disjoint training set, STEMMA-Instruct. Fine-tuning two LALMs on STEMMA-Instruct improves multi-audio reasoning, with the largest gains on structural relations directly determined by the catalog, while preserving single-audio music understanding.

Comments5 pages, 1 figure, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑