CLEAR:通过跨ASR不一致性进行在线语音内容泄露估计
CLEAR: Online Speech Content Leakage Estimation through Cross-ASR Disagreement
浏览论文内容
中文总结 AI 辅助
CLEAR提出一种无参考的在线语音内容泄露估计方法,利用异构ASR系统的不一致性在运行时评估隐私保护效果,实现动态调整隐私强度并保留声学效用。
中文摘要 AI 辅助
信号级语音隐私机制在抑制语言内容的同时,保留下游传感应用所需的声学信息。然而,其隐私设置通常在离线状态下评估/选择,并在部署期间保持固定,尽管语音内容泄露在不同话语和说话者之间可能差异很大。在运行时调整隐私保护需要估计有多少语音仍可恢复,但诸如WER或PER等传统指标需要真实转录文本,因此无法在线计算。我们提出CLEAR,一种利用异构ASR系统之间的不一致性在运行时估计语音内容泄露的无参考方法。我们的关键见解是,独立训练的ASR在语言内容可恢复时表现出一致行为,而随着隐私变换模糊语音,它们越来越不一致。使用可配置的语音抑制机制,我们表明跨ASR不一致性在隐私操作点上紧密跟踪基于转录的泄露,达到0.8的相关性。我们进一步刻画了异构ASR子集的延迟-准确性权衡,并利用它们的假设来识别可能暴露的单词。这些能力使隐私能够被视为运行时属性而非固定配置:CLEAR可以向用户传达残余语音暴露,并提供反馈以动态调整隐私强度,同时保留下游传感任务的声学效用。
英文摘要
Signal-level speech privacy mechanisms suppress linguistic content while preserving acoustic information needed by downstream sensing applications. However, their privacy settings are typically evaluated/selected offline and remain fixed during deployment, even though speech-content leakage can vary substantially across utterances and speakers. Adapting privacy protection at runtime requires estimating how much speech remains recoverable, but conventional measures such as WER or PER require ground-truth transcripts and therefore cannot be computed online. We present CLEAR, a reference-free approach for estimating speech-content leakage at runtime using disagreement among heterogeneous ASR systems. Our key insight is that independently trained ASRs exhibit consistent behavior when linguistic content remains recoverable and increasingly disagree as privacy transformations obscure speech. Using configurable speech-suppression mechanism, we show that cross-ASR disagreement closely tracks transcript-grounded leakage across privacy operating points, achieving a correlation of 0.8. We further characterize the latency-accuracy trade-off of heterogeneous ASR subsets and use their hypotheses to identify potentially exposed words. These capabilities enable privacy to be treated as a runtime property rather than a fixed configuration: CLEAR can communicate residual speech exposure to users and provide feedback for dynamically adjusting privacy aggressiveness while retaining acoustic utility for downstream sensing tasks.
发表机构
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。