发表机构
independent researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究配音演员语音克隆归因问题,指出自然防御在声音嵌入空间拥挤时失败,存在错误识别下限和错误归因,通用编码器不公平,提出需扩展防欺骗感知说话者验证到开放集1:N来实现稳健归因。
AI 中文摘要
配音演员的声音是其资产,而人工智能克隆直接威胁到这一点。自然防御会标记出嵌入相似度与可疑录音超过阈值的已注册演员。但我们发现它在最需要的地方失败了:训练过的声音挤满了嵌入空间,且每个演员有多种风格。在1168名日本配音演员(56568个片段,约63小时)上,错误识别下限在校准、分数归一化和判别性重新排序(线性和非线性,包括PLDA)后仍然存在:残差是嵌入几何的限制,而非我们评估的后端的限制。最佳集成仍有约2.6%的闭集错误识别,比匹配的对照组高出几倍;会话不相交、重新排序仅将下限降至13.0%。同样的拥挤导致错误归因:在通用英语编码器上,未注册人员的大约一半克隆错误地指控已注册演员,而在相同阈值下,已注册目标的Seed-VC克隆有32%被遗漏;一个操作点将两者联系起来,且没有一个能同时避免两者。域匹配、配音演员训练的编码器大幅减轻了问题(四倍性别差距消失;错误错误归因降至1.5 - 10%),但未消除下限。对照组(编解码器、通道、声码器、内容)支持将遗漏率视为真实与合成协变量转移,而非缺少说话者信息。因此,固定阈值克隆归因在此不可靠,在通用编码器上不公平。稳健归因必须将防欺骗感知说话者验证扩展到开放集1:N(防欺骗门、域匹配编码器、每个说话者校准、弃权选项),即便如此也仅支持检测,而非自主执行。
英文摘要
A voice actor's voice is their asset, and AI cloning directly threatens it. The natural defense flags the enrolled actor whose embedding similarity to a suspect recording crosses a threshold. We show it fails where it is most needed: trained voices crowd the embedding space, and each actor performs many styles. On 1,168 Japanese voice actors (56,568 segments, ~63 h), a misidentification floor survives calibration, score normalization, and discriminative re-ranking (linear and nonlinear, including PLDA): the residual is a limit of the embedding geometry, not of the back-ends we evaluate. The best ensemble still leaves ~2.6% closed-set misidentification, several-fold above matched controls; session-disjoint, re-ranking lowers the floor only to 13.0%. The same crowding drives false attribution: on a generic English encoder, roughly half the clones of non-enrolled people falsely accuse an enrolled actor, while -- by a separate real-vs-synthetic shift -- 32% of Seed-VC clones of enrolled targets are missed at the same threshold; one operating point couples the two, and none escapes both. A domain-matched, voice-actor-trained encoder mitigates substantially (a four-fold gender gap vanishes; wrongful misattribution falls to 1.5-10%), but does not remove the floor. Controls (codec, channel, vocoder, content) support reading the miss rate as a real-versus-synthetic covariate shift, not missing speaker information. Fixed-threshold clone attribution is thus unreliable here, and on a generic encoder unfair. Robust attribution must extend spoofing-aware speaker verification to open-set 1:N (anti-spoofing gate, domain-matched encoder, per-speaker calibration, abstain option), and even then supports detection, not autonomous enforcement.
Comments24 pages, 6 figures, 8 tables. Submitted to IEEE Access. v2: adds related work on two concurrent ASJ Spring 2026 studies of acting voices (Yamamoto et al.; Hayashi et al.) with corresponding discussion and limitations updates; results unchanged