发表机构
Tsinghua University; The Chinese University of Hong Kong; SenseTime(清华大学; 香港中文大学; 商汤科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究流式目标说话人提取中的质量-可懂度权衡问题,提出扩大Conformer卷积核及基于WavLM的DPO微调策略,实现了相对可懂度10.9%的提升,同时音频质量和说话人相似度也有改善。
AI 中文摘要
用于目标说话人提取(TSE)的生成式流式模型通常存在质量-可懂度权衡问题,即单纯优化感知音频质量会降低语音可懂度,反之亦然。我们发现这种权衡并非源于流式架构的限制,而是优化锚点选择不当。直接针对音频质量指标进行优化会引发灾难性的奖励作弊,关键发音和可懂度的内容会被系统删除以最大化代理分数。为打破这一瓶颈,我们提出两项互补改进:扩大的Conformer卷积核用于更丰富的局部频谱-时间建模,以及基于WavLM的直接偏好优化(DPO)微调策略。DPO偏好对通过WavLM余弦相似度排序,其深度声学特征编码了语音结构和说话人身份,提供了抗作弊的优化锚点。在560毫秒的流式块大小下,该方法实现了10.9%的相对可懂度提升(字错误率从0.138降至0.123),同时音频质量和说话人相似度略有提升。
英文摘要
Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely. We reveal that this trade-off arises not from the constraints of streaming architectures, but from an inappropriate choice of optimization anchor. Directly optimizing against audio quality metrics induces catastrophic reward hacking, where content critical to pronunciation and intelligibility is systematically erased to maximize a proxy score. To break this bottleneck, we propose two complementary improvements: an enlarged Conformer convolution kernel for richer local spectro-temporal modeling, and WavLM-anchored Direct Preference Optimization (DPO) fine-tuning strategy. DPO preference pairs are ranked by WavLM cosine similarity, a deep acoustic feature encoding both phonetic structure and speaker identity, providing an optimization anchor that resists hacking. Under a 560 ms streaming chunk size, the proposed method achieves a 10.9% relative intelligibility improvement (word error rate: 0.138 to 0.123), with marginal simultaneous gains in audio quality and speaker similarity.