发表机构
University at Buffalo; Southwest University of Political Science and Law; Australian National University(布法罗大学; 西南政法大学; 澳大利亚国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探讨“声纹谬误”,指出语音并非稳定独特的生物识别印记,通过回顾相关研究,建议用考虑变异性与不确定性的概率框架解读语音证据。
AI 中文摘要
近年来,“声纹”一词重新受到关注,尤其在技术应用和政策制定领域,人们常假设人的语音构成一种稳定且独特的生物识别痕迹,类似于指纹。然而,自声纹概念提出以来的数十年间,法医语音专家多次批评并拒绝了这一构想。尽管语音无疑包含与说话人相关的信息,但这种简化的构想掩盖了语音高度动态且依赖语境的本质。本文通过回顾声纹识别的历史发展、人类语音变异性的相关证据、法医语音比对的进展、人类与自动说话人识别的研究,以及深度伪造语音对说话人身份构成的最新挑战,重新审视声纹谬误,并重新考量何种内容可作为说话人身份的证据。我们指出,声纹隐喻及其隐含的意义在科学上具有误导性,因为它们将说话人信息的概率来源转化为想象中稳定的身份对象。为避免将语音视为印记般的痕迹,我们建议通过经验证和校准的概率框架来解读语音证据,该框架需明确考虑变异性、不确定性及其他可能的解释。
英文摘要
In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the assumption that a person's voice constitutes a stable and unique biometric trace analogous to a fingerprint. Yet this conception has been repeatedly criticized and rejected by forensic voice experts throughout the decades since its introduction. Although voices undoubtedly contain speaker-related information, this simplified conception obscures the highly dynamic and context-dependent nature of speech. This article revisits the voiceprint fallacy and reconsiders what can count as evidence of speaker identity by reviewing the historical development of voiceprint identification, evidence on human voice variability, developments in forensic voice comparison, research on human and automatic speaker recognition, and the recent challenge posed by deepfake speech to speaker identity. We point out that the voiceprint metaphor and its underlying implications are scientifically misleading because they transform a probabilistic source of speaker information into an imagined stable object of identity. We argue that speaker identity assessment does not require, and current evidence does not support, the existence of a stable and individually unique voiceprint. For speaker recognition and voice biometrics, this distinction motivates interpreting learned speaker representations with respect to the conditions under which they are trained and evaluated, and explicitly assessing their robustness to relevant sources of within-speaker variability, domain mismatch, and synthetic manipulation.