角色身份并非说话人身份:KyaraBench 与 KyaraEmbed 用于角色验证
Character Identity is not Speaker Identity: KyaraBench and KyaraEmbed for Character Verification
浏览论文内容
中文总结 AI 辅助
针对角色身份与说话人身份不同的问题,提出 KyaraBench 基准和 KyaraEmbed 编码器,直接验证角色声音,并在所有角色特定条件下取得最优性能。
中文摘要 AI 辅助
配音角色在更换声优后仍保持其身份,因此角色身份与说话人身份是同一录音的两个不同属性,而说话人验证仅衡量后者。为弥补这一差距,我们提出了 KyaraBench,一个直接对角色声音而非说话人声音进行评分的基准,该基准基于一个配音动漫语料库中经人工审核的 85 个身份构建。该基准设置了两个具有挑战性的条件:一是在固定角色下更换表演者,二是在角色变化下固定表演者。我们邀请了 78 名参与者进行听力研究,为跨表演者验证和同演员区分提供了人类参考分数。说话人验证基线在这些角色特定条件下表现出更高的错误率。随后,我们训练了 KyaraEmbed,一个紧凑的编码器,利用多语言角色监督、同演员负样本以及语言对齐项。该模型在主比较中所有角色特定条件下均取得了最佳性能。我们公开了该基准、协议和编码器。
英文摘要
A dubbed character keeps its identity while the voice actor changes, so character identity and speaker identity are distinct properties of one recording, yet speaker verification measures only the latter. To address this gap, we propose KyaraBench, a benchmark that scores character voice directly instead of speaker voice, built from 85 human-audited identities in a dubbed anime corpus. It poses two challenging conditions: one that swaps the performer under a fixed character, and one that fixes the performer under changing characters. Listening studies with 78 participants provide human reference scores for cross-performer verification and same-actor discrimination. Speaker-verification baselines show increased errors under these character-specific conditions. We then train KyaraEmbed, a compact encoder using multilingual character supervision, same-actor negatives, and a language-alignment term. The model achieves the best performance on all character-specific conditions in the main comparison. We release the benchmark, protocol, and encoder publicly.
发表机构
- Spellbrush
机构由 AI 辅助整理,请以论文原文为准。