发表机构
Inventec Corporation; National Taiwan University(英业达股份有限公司; 台湾大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对指令TTS中的性别偏差,提出无训练模型自适应转向方法,通过组偏差向量与强度搜索校准,在四个模型上显著降低校准误差,同时保持语音质量。
AI 中文摘要
指令文本到语音(ITTS)系统会从描述性风格提示(如职业或人物角色)中系统性地编码声学性别偏差,即使未指定人口统计属性也是如此。在不进行模型重训练或覆盖用户明确提示的情况下,跨异构架构校准合成语音的隐含性别分布是一个开放性问题。在本工作中,我们提出模型自适应转向,一种无训练的偏差校准方法,该方法使用组偏差向量结合开发集上的从粗到细的强度搜索来引导编码器后条件表示。一个确定性词汇门在检测到明确性别关键词时绕过干预,保留预期的提示语义。在包含12,800个提示的保留基准上,对四个ITTS模型(三种架构)进行评估,所提方法将聚合校准误差从11.5–28.1个百分点降低至0.8–5.9个百分点(例如,将Parler-TTS Mini的女性比例从78.1%和PromptTTS++的27.3%调整至50.8–52.4%)。在固定操作点下,输出速率在锚定集大小范围内保持在0.9个百分点以内,而UTMOS最多下降0.08,WER最多退化3.3个百分点。
英文摘要
Instruction text-to-speech (ITTS) systems systematically encode acoustic gender skews from descriptive style prompts, such as occupations or personas, even when demographic attributes are left unspecified. Calibrating the implicit gender distribution of the synthesized voices, across heterogeneous architectures and without model retraining or overriding explicit user prompts, is an open problem. In this work, we propose \textit{model-adaptive steering}, a training-free bias calibration method that steers post-encoder conditioning representations using a group deviation vector paired with a coarse-to-fine strength search on a development set. A deterministic lexical gate bypasses intervention whenever explicit gender keywords are detected, preserving intended prompt semantics. Evaluated on a 12,800-prompt held-out benchmark across four ITTS models (three architectures), the proposed method reduces aggregate calibration error from 11.5--28.1 to 0.8--5.9 percentage points (e.g., shifting Parler-TTS Mini from 78.1\% and PromptTTS++ from 27.3\% female to 50.8--52.4\%). Output rates stay within 0.9 pp across anchor set sizes at fixed operating points, while UTMOS decreases by at most 0.08 and WER degrades by at most 3.3 pp.
Comments5 pages, 1 figure, 3 tables, submitted to ICASSP 2027