arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17098cs.SDcs.CR

通过有针对性的对抗性混淆和表征学习实现基于语音的多级隐私保护痴呆症检测

Multi-Level Privacy-Preserving Dementia Detection from Speech via Targeted Adversarial Obfuscation and Representation Learning

Henriette Flore Kenne, Raphael Anaadumba, Mohammad Arif Ul Alam

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对痴呆症检测语音记录的隐私问题,提出多级框架,信号层用CSA、特征层用带MI引导噪声注入的GRL,在痴呆症银行匹兹堡语料库上评估,兼顾隐私保护与痴呆症分类性能。

中文摘要 AI 辅助

用于痴呆症检测的语音记录会暴露说话者身份,引发隐私问题。现有方法通常只解决单一威胁,无法解决隐私与效用的权衡。我们提出一个多级框架来消除两种不同的窃听向量。在信号层面,累积信号攻击(CSA)将扰动集中在关键词对齐区域以最大化转录错误(字错误率WER = 1.00),同时保留重要的韵律生物标志物。在特征层面,带有互信息(MI)引导噪声注入的梯度反转层(GRL)抑制说话者判别维度,同时保留与痴呆症相关的诊断结构。在痴呆症银行匹兹堡语料库上评估,我们的框架实现了接近随机的说话者识别(等错误率EER = 0.59,F1 = 0.003),同时保持了强大的痴呆症分类性能(F1 = 0.78,AUC = 0.86)。

英文摘要

Speech recordings used for dementia detection inherently expose speaker identity, raising critical privacy concerns. Existing methods typically address only singular threats and fail to resolve the privacy--utility trade-off. We propose a multi-level framework designed to neutralize two distinct eavesdropping vectors. At the signal level, a Cumulative Signal Attack (CSA) concentrates perturbations in keyword-aligned regions to maximize transcription error (Word Error Rate WER = 1.00) while preserving vital prosodic biomarkers. At the feature level, a Gradient Reversal Layer (GRL) with Mutual Information (MI)-guided noise injection suppresses speaker-discriminative dimensions while retaining dementia-relevant diagnostic structure. Evaluated on the DementiaBank Pitt Corpus, our framework achieves near-chance speaker identification (Equal Error Rate EER = 0.59, F1 = 0.003) while maintaining strong dementia classification performance (F1 = 0.78, AUC = 0.86).

发表机构

  • University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑