Decaf:一种利用说话人解耦和规范语音转换的隐私保护语音编解码器
Decaf: A privacy preserving speech codec using speaker disentanglement and canonical voice conversion
浏览论文内容
中文总结 AI 辅助
DECAF是一种隐私保护语音编解码器,通过说话人解耦和规范语音转换在0.5kbps下混淆声音,保持ASR性能,EER达43.5%,词错误率相对降低33.2%。
中文摘要 AI 辅助
我们提出DECAF,一种隐私保护的神经语音编解码器,受“脱咖啡因”启发,在极低比特率下混淆说话人的声音,同时保留语言内容并维持自动语音识别(ASR)性能。在发送端,语音被编码为与说话人无关的内容嵌入,这些嵌入通过残差向量量化进行压缩,并在不携带任何说话人相关信息的情况下传输。在接收端,使用端点间预先共享的规范说话人嵌入进行波形重建,从而实现对说话人声音的确定性和一致性混淆。所提出的框架利用应用于自监督表示的信息瓶颈,以及独立的说话人嵌入分支,以实现有效的说话人内容解耦。我们进一步引入基于CTC的辅助目标,鼓励内容表示与下游ASR任务良好对齐。我们表明,DECAF在0.5 kbps的比特率下运行,对说话人验证系统实现了高达43.5%的等错误率(EER),同时保持具有竞争力的ASR性能,与最先进的方法相比,词错误率相对降低了33.2%。
英文摘要
We present DECAF, a privacy preserving neural speech codec that obfuscates a speaker's voice while preserving linguistic content while maintaining automatic speech recognition (ASR) performance at very low bitrates, inspired by decaffeination. At the transmitter end, speech is encoded into speaker independent content embeddings, which are compressed using residual vector quantization and transmitted without any speaker related information. At the receiver, a canonical speaker embedding, shared a priori between endpoints, is used for waveform reconstruction, enabling deterministic and consistent obfuscation of a speaker's voice. The proposed framework leverages an information bottleneck applied to self supervised representations, along with a separate speaker embedding branch, to achieve effective speaker content disentanglement. We further incorporate a CTC-based auxiliary objective, encouraging content representations that are well aligned with downstream ASR tasks. We show that DECAF operating at a bit rate of 0.5 kbps achieves an Equal Error Rate (EER) of up to 43.5% for a speaker verification system, while maintaining competitive ASR performance, yielding a relative reduction in word error rate of 33.2% compared to a state of the art method.
发表机构
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。