文本生成音乐模型真的遵循指令吗?对调式与节拍分组的反事实评估
Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping
浏览论文内容
中文总结 AI 辅助
本研究提出匹配反事实评估方法,评估文本生成音乐模型对调式、节拍分组指令的遵循情况,发现不同模型在调式和节拍控制上存在差异,该方法可修正相关经验结论。
中文摘要 AI 辅助
提示属性一致性被广泛用作文本生成音乐可控性的证据,但请求的属性可能仅仅因为它在模型输出分布中已经常见而出现。我们引入一种匹配反事实评估,将目标属性的出现与指令可归因的控制分离开来。每个系列包含一个省略了评分属性的中性输入,以及两个其他方面匹配但交换了请求目标的输入,所有三个输入都通过具有共享种子的冻结原生接口适配器渲染。将该设计应用于三个开放系统的全局调式和节拍分组时,它改变了经验结论:ACE-Step 1.5和Stable Audio 3 Medium表现出显著的调式控制,而LeVo2则没有。对于节拍分组,相同模型会转向罕见的三拍目标,但高四拍一致性很大程度上继承自中性输出:Stable Audio 3在0.97的中性案例中产生四拍分组,而在明确的四拍处理下仅为0.56。非属性安慰剂、外部识别器验证、盲专家标注和多种子哨兵支持该归因。当目标具有不等的输出先验时,一致性描述模型产生的内容,而匹配的中性和目标交换对比则测试指令是否改变了它。
英文摘要
Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's output distribution. We introduce a matched counterfactual evaluation that separates target occurrence from instruction-attributable control. Each family contains a neutral input that omits the scored attribute and two otherwise matched inputs that swap the requested target. All three are rendered through frozen native-interface adapters with a shared seed. Applied to global key and beat grouping in three open systems, this design changes the empirical conclusion. ACE-Step 1.5 and Stable Audio 3 Medium exhibit substantial key control, whereas LeVo2 does not. For beat grouping, the same models redirect toward the rare three-beat target, but high four-beat agreement is largely inherited from neutral outputs: Stable Audio 3 produces four-beat grouping in 0.97 of neutral cases but only 0.56 under its explicit four-beat treatment. Off-attribute placebos, external recognizer validation, blind expert annotation, and multi-seed sentinels support the attribution. When targets have unequal output priors, agreement describes what a model produced, while matched neutral and target-swap contrasts test whether the instruction changed it.