发表机构
National Taiwan University; Amazon(国立台湾大学; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究文本到音频指令跟随问题,提出用音频感知大语言模型作细粒度评判器的框架,经验证后用其反馈构建偏好对优化,引入S3Bench基准,实验证明该方法能提升事件完整性、时间排序和指令跟随准确性,且保持音频质量。
AI 中文摘要
近期文本到音频模型能生成高质量音频,但在处理涉及多个声音事件和时间顺序的指令时常常失败。现有评估和训练信号主要强调全局相似性或感知质量,对指令级正确性监督有限。我们提出一个指令级框架,利用音频感知大语言模型作为细粒度评判器来验证生成音频中目标事件的存在和时间关系。在基准测试上验证大语言模型的判断并经人工验证后,利用其反馈构建偏好对进行直接偏好优化。我们还引入了S3Bench,一个用于评估多事件时间指令跟随的叙事基准。实验表明,我们的方法在保持音频质量的同时,提高了现有基准和S3Bench上的事件完整性、时间排序和联合指令跟随准确性。
英文摘要
Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchmarks and S3Bench, while maintaining audio quality.
CommentsAccepted to the Long Paper Track at Interspeech 2026. Project Website: https://kuan2jiu99.github.io/allm-feedback-tta