AI 中文总结
本研究提出一种轻量无需训练的开放词汇概念发现方法,用于可控音乐检索,可恢复与概念音频示例对齐更紧密的稀疏特征,实现更优的编辑强度与内容保留权衡。
AI 中文摘要
可控音乐检索允许用户在保留原始种子查询其他语义内容的同时,查找更具氛围感、失真度更低或不含吉他等特征的音乐。稀疏自编码器(SAEs)是实现此类概念级控制的有前景接口,但仍存在一个关键问题:给定自由形式的文本概念,应编辑哪些稀疏特征?在共享多模态嵌入空间中,标准归因方法通常会选择与概念措辞匹配但与表达该概念的音频示例不匹配的神经元,这会导致编辑效果弱或不稳定:当概念分布在多个神经元上时,会遗漏相关特征;而其他特征则因文本对齐而非音频侧结构被选中。我们提出一种轻量、无需训练的方法,可恢复一组稀疏音频特征,其解码表示能重构目标概念,同时与音频空间几何保持一致。该方法将概念归因重新定义为稀疏逆问题,而非文本侧神经元排序启发式方法,且无需配对音频-文本监督或SAE再训练。我们在可控音乐检索中评估该方法,结果显示,所恢复的支持项与承载概念的音频示例对齐更紧密,且在编辑强度与内容保留之间实现了比对齐基线更优的权衡,能实现更精确的概念放大与抑制,同时降低保留指标上的漂移。
英文摘要
Controllable music retrieval lets users find music that is, for example, more ambient, less distorted, or without guitar while preserving the other semantic content of an original seed query. Sparse autoencoders (SAEs) are a promising interface for this kind of concept-level control, but a key problem remains: given a free-form text concept, which sparse features should be edited? In shared multimodal embedding spaces, standard attribution methods often select neurons that match the concept's wording but not the audio examples that express it. This leads to weak or unstable edits: relevant features are missed when concepts are distributed across neurons, while others are selected due to text alignment rather than audio-side structure. We address this with a lightweight, training-free method that recovers a sparse set of audio features whose decoded representation reconstructs the target concept while remaining consistent with audio-space geometry. This reframes concept attribution as a sparse inversion problem rather than a text-side neuron-ranking heuristic. The method requires neither paired audio-text supervision nor SAE retraining. We evaluate this approach in steerable music retrieval and show that the recovered supports align more closely with concept-bearing audio examples and achieve a stronger trade-off between edit strength and preservation than alignment baselines, enabling more precise concept amplification and suppression with reduced drift on preservation metrics.
CommentsAccepted to ISMIR 2026