发表机构
University of California, Los Angeles; University of Tennessee, Knoxville(加州大学洛杉矶分校; 田纳西大学诺克斯维尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对微调TTS模型提出首个黑盒成员推断框架,通过优化查询生成与表示工程,在三类模型上实现高隐私泄露,揭示了TTS模型的严重隐私风险。
AI 中文摘要
文本转语音(TTS)基础模型越来越多地在私有数据集上进行微调,以合成高度个性化的语音,这会带来严重的隐私风险,暴露生物识别身份和敏感语音内容。现有的黑盒成员推断攻击(MIAs)遵循查询生成和表示工程的两阶段流程,当应用于TTS时,这两个阶段都面临独特的挑战。对于查询生成,合成文本和参考语音的双重条件创建了一个庞大且未被充分探索的查询设计空间,尚无确定的标准来识别有效查询。对于表示工程,语音的多级特征和时间变异性使得低级表示和直接比较不足以捕捉成员信号。为应对这些挑战,我们提出了第一个专门针对TTS模型的黑盒MIA框架,覆盖说话者和记录两个层面。在查询生成方面,我们刻画了可行的查询空间,建立了可评分程度和记忆诱导两个标准来评估五个代表性查询,确定朗诵是最强的查询。在表示工程方面,我们从嵌入模型获取多级语音表示,并将生成音频与目标音频进行时间对齐以进行细粒度比较。在两个基准数据集(VCTK和British Dialect)上微调的三个最先进TTS模型(CosyVoice2、F5-TTS和XTTS-v2)的评估显示存在严重隐私泄露:说话者级AUC保持在0.80以上,在最强设置下接近1.0;记录级AUC范围为0.80至0.90,即使在成员和非成员属于同一说话者的具有挑战性场景中也依然有效。我们进一步确定了与记忆脆弱性不成比例相关的语音特征。
英文摘要
Text-to-Speech (TTS) foundation models are increasingly fine-tuned on private datasets to synthesize highly personalized voices, introducing severe privacy risks by exposing both biometric identities and sensitive speech content. Existing black-box membership inference attacks (MIAs) follow a two-stage pipeline of query generation and representation engineering, both of which face unique challenges when adapted to TTS. For query generation, dual conditioning on synthesis text and reference speech creates a large and underexplored query design space with no established criterion for identifying an effective query. For representation engineering, the multi-level speech characteristics and temporal variability of speech make low-level representations and direct comparisons inadequate for capturing membership signals. To address these challenges, we present the first black-box MIA framework explicitly tailored to TTS models at both the speaker and record levels. For query generation, we characterize the feasible query space and establish two criteria, scorable extent and memorization elicitation, for evaluating five representative queries, identifying recitation as the strongest. For representation engineering, we obtain multi-level speech representations from embedding models and temporally align the generated and target audio for fine-grained comparison. Evaluations across three state-of-the-art TTS models (CosyVoice2, F5-TTS, and XTTS-v2) fine-tuned on two benchmark datasets (VCTK and British Dialect) reveal severe privacy leakage: speaker-level AUC remains above 0.80 and approaches 1.0 in the strongest settings, while record-level AUC ranges from 0.80 to 0.90 and remains effective even in challenging scenarios where both members and non-members are of the same speakers. We further identify speech characteristics associated with disproportionate vulnerability to memorization.
Comments18 pages