arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

“我见即我所闻”:跨不同听力能力的深度伪造检测

"What I See is What I Hear": Deepfake Detection Across Diverse Hearing Abilities

Magdalena Pasternak, Malvika Jadhav, Palavi V. Bhole, Aviva Smith, Elaina Trapatsos, Vincent Bindschaedler, Roshan Peiris, Ersin Uzun, Patrick Traynor, Matthew Wright, Kevin R. B. Butler

arXiv 2609.28659首次发表:更新:

发表机构

University of Florida; Rochester Institute of Technology(佛罗里达大学; 罗切斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过80名不同听力能力参与者的实验,发现聋哑及重听人群在深度伪造检测中准确率更低,主要因将真实片段误判为伪造,凸显了针对不同听力能力群体的定制化防御需求。

AI 中文摘要

音视频深度伪造的泛滥降低了欺诈、冒充和虚假信息的成本,但其成功最终取决于人类的感知。检测需要整合听觉和视觉线索,然而安全和隐私研究在很大程度上忽视了聋哑及重听(DHH)人群。我们通过一项面对面的混合方法研究来弥补这一空白,该研究涉及80名参与者:31名听力正常者(HPs)、15名重听(HoH)参与者、17名聋哑参与者和17名人工耳蜗(CI)使用者。每位参与者对30个片段的真实性进行了判断,其中操纵涵盖了文本转语音、语音转换、唇形同步或换脸。DHH参与者整体上比HP的准确率更低(76.4%对88.0%,p<.001),主要原因是他们更频繁地将真实片段分类为被操纵的(假阳性率:29.7%对11.2%)。差异在很大程度上取决于被操纵的通道。对于仅音频操纵,HoH参与者与HP的表现相当(90.0%对90.3%),其次是CI使用者(79.4%)和聋哑参与者(41.2%)。当片段包含音视频操纵时,准确率集中在84%至87%之间,尽管性能仍因操纵方法而异。我们的工作系统地描述了深度伪造如何影响DHH人群,强调了音视频操纵可能对不同听力能力群体构成的不对称风险,以及支持所有用户的可访问、量身定制的防御措施的必要性。

英文摘要

The proliferation of audiovisual deepfakes has lowered the cost of fraud, impersonation, and misinformation, but their success ultimately depends on human perception. Detection requires integrating auditory and visual cues, yet security and privacy research has largely overlooked d/Deaf and hard-of-hearing (DHH) populations. We address this gap with an in-person, mixed-methods study of 80 participants: 31 hearing persons (HPs), 15 hard-of-hearing (HoH) participants, 17 d/Deaf participants, and 17 cochlear implant (CI) users. Each participant judged the authenticity of 30 clips, where manipulations spanned text-to-speech, voice conversion, lip-sync, or face-swap. DHH participants were less accurate than HPs overall (76.4% vs. 88.0%, p<.001), primarily because they more often classified authentic clips as manipulated (FPR: 29.7% vs. 11.2%). Differences depended strongly on the manipulated channel. For audio-only manipulations, HoH participants matched HPs (90.0% vs. 90.3%), followed by CI users (79.4%) and d/Deaf participants (41.2%). When clips contained an audiovisual manipulation, accuracy clustered between 84% and 87%, although performance still varied by manipulation method. Our work systematically characterizes how deepfakes affect DHH populations, highlighting the asymmetric risks audiovisual manipulations may pose to groups with different hearing abilities and the need for accessible, tailored defenses that support all users.

CommentsProceedings of the Network and Distributed System Security (NDSS) Symposium 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑