发表机构
KTH Royal Institute of Technology(瑞典皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过微调潜在音频扩散模型适配器,并辅以专家盲听评估,成功将模型适配至历史古琴录音,实现了高质量的古琴风格音乐续奏与场景生成。
AI 中文摘要
我们报告了一项小数据案例研究,旨在将预训练的潜在音频扩散模型适配至古琴(七弦中国齐特琴),目标是为一个不间断播放古琴风格音乐的“AI电台”提供支持。我们从历史录音库中精选出412首独奏曲目(总计42.5小时,涉及61位演奏者),并按乐曲进行划分。我们使用Stable Audio 3 Medium的连续潜在表示,微调了一个秩为16的DoRA适配器,其掩码加权偏向于续奏,标题结合了每首乐曲的研究笔记、情绪标签以及自动估计的五声音阶模式,并采用随机长度裁剪,在验证集损失停止改善时停止训练。七项由一位专家听众进行的盲听研究指导了每一项决策。最终适配器在续奏未见乐曲方面获得最高评分(3.9/5,而此前最佳适配器为3.4/5),在根据自由撰写的场景描述生成音乐方面获得4.6/5的评分。负面结果同样具有信息量:一个在相同潜在表示上从头训练的自回归模型未能产生可辨识的音色,一种信号级摩擦噪声测量与听众的抱怨呈错误方向相关,而五声音阶拟合度量总体上与评分相关,但在候选组内几乎无区分度。链式续奏用于长时间播放时,每个生成片段均暴露出静音尾部以及响度反馈回路问题,两者均有简单修复方案。由于仅有一位听众且每种条件下最多十个片段,没有配对差异具有统计显著性;我们呈现了一份探索性记录,说明哪些方法有效、哪些无效及其原因。
英文摘要
We report a small-data case study in adapting a pretrained latent audio diffusion model to the guqin, the seven-string Chinese zither, aiming at an "AI radio" that plays guqin-style music without end. From a library of historical recordings we curate 412 solo performances (42.5 h, 61 performers) and split them by composition. We fine-tune a rank-16 DoRA adapter on Stable Audio 3 Medium using its continuous latents, masking weighted towards continuation, captions that combine researched notes on each piece with mood tags and an automatically estimated pentatonic mode, and random-length crops, stopping when held-out loss stops improving. Seven blind listening studies by one expert listener guided every decision. The final adapter was rated highest for continuing unseen pieces (3.9/5, against 3.4 for the best earlier adapter) and 4.6/5 for generating from free-written scene descriptions. Negative results are equally informative: a from-scratch autoregressive model over the same latents produced no recognisable timbre, a signal-level friction-noise measure correlated with the listener's complaints in the wrong direction, and a pentatonic-fit measure tracked ratings overall but barely within a group of candidates. Chaining continuations for long playback exposed a silent tail on every generated clip and a loudness feedback loop, both with simple fixes. With one listener and at most ten clips per condition, no paired difference is statistically significant; we present an exploratory record of what helped, what did not, and why.