发表机构
The Chinese University of Hong Kong, Shenzhen; Microsoft Research(香港中文大学(深圳); 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AuraSE通过多模态流匹配与推理策略优化,减少生成式语音增强中的幻觉,在合成和真实测试集上取得领先性能。
AI 中文摘要
生成式语音增强模型能够比传统的判别式方法产生更干净、更自然的语音,但即使在有转录条件的情况下,也可能通过改变语音内容或说话人身份而产生幻觉。我们提出了AuraSE,一个流匹配框架,通过互补的模态和推理设计来解决幻觉问题。首先,双流到单流的多模态扩散Transformer(MMDiT)允许转录和声学表示相互作用,同时为退化输入保留专用路径。其次,我们发现由引导尺度、采样温度和步数控制的最佳解码器配置在不同话语间差异很大。这一观察促使我们提出推理策略优化(IPO),一种在线的、同策略的偏好优化方法。IPO从当前模型在不同推理配置下生成多个候选,使用多目标奖励对它们进行排序,并从它们的相对偏好中学习。AuraSE-IPO在合成测试集的12项指标中排名第一,并在真实DNS盲测集上获得评估系统中最高的DNSMOS和盲听分数。在部署时,它使用固定的10步ODE解码器,无需分类器自由引导(CFG)或逐话语配置搜索。
英文摘要
Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framework that addresses hallucination through complementary modality and inference designs. First, a double-stream-to-single-stream multimodal Diffusion Transformer (MMDiT) allows transcript and acoustic representations to interact while preserving a dedicated pathway for the degraded input. Second, we find that the best decoder configuration, governed by guidance scale, sampling temperature, and step count, varies substantially across utterances. This observation motivates Inference Policy Optimization (IPO), an online, on-policy preference optimization method. IPO generates multiple candidates from the current model under different inference configurations, ranks them with a multi-objective reward, and learns from their relative preferences. AuraSE-IPO ranks first on 11 of 12 metrics across the synthetic test sets and obtains the highest DNSMOS and blind-listening scores among the evaluated systems on the real DNS blind test set. At deployment, it uses a fixed $10$-step ODE decoder without classifier-free guidance (CFG) or per-utterance configuration search.