arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

去除语音,保留活动:面向辅助生活声学感知的隐私防火墙

Removing Speech, Keeping Activities: A Privacy Firewall for Acoustic Sensing in Assisted Living

Pavlos Nicolaou, Christos Efstratiou

arXiv 2609.02376首次发表:更新:

发表机构

KIOS Research and Innovation Center of Excellence, University of Cyprus; School of Computing, University of Kent(塞浦路斯大学KIOS卓越研究与创新中心; 肯特大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于U-Net编解码器的隐私防火墙流水线,可去除环境音频中的语音并保留活动相关环境音,在ESC-50、SINS及真实家庭录音上验证了其隐私保护与活动识别性能,解决老年护理声学感知的隐私障碍。

AI 中文摘要

声学感知为监测老年人日常活动提供了一种有前景的非侵入式方法,但语音隐私问题仍是其实际部署的关键障碍。我们提出一种基于U-Net编解码器的隐私防火墙流水线,该模型完全在合成数据上训练,可从环境音频中去除语音,同时保留指示日常活动的环境声音。活动识别采用VGGish迁移学习结合SVM分类器实现。在ESC-50和SINS数据集上,针对多种语音内容水平的评估显示,所提模型在所有测试条件下均将残留语音降至0%(Silero语音活动检测VAD可检测的语音),在100%语音水平的ESC-50数据集上,其性能优于Facebook Denoiser(残留6.55%)、SepFormer(残留36.34%)和ConvTasNet(残留47.21%)。在40%语音水平的ESC-50数据集上,去除语音后分类性能恢复至85%的精确率和85%的召回率,而去除前为81%/75%,无语音的基线为84%/83%。对通过AudioHive应用收集的真实世界参与者家庭录音的评估显示,处理后VAD可检测的语音为0%,同时保持76%的精确率和召回率。该流水线可在不牺牲活动识别性能的情况下实现隐私保护的声学感知,解决了环境监测在老年护理中应用的关键障碍。

英文摘要

Acoustic sensing offers a promising non-intrusive approach for monitoring daily activities of older adults, yet speech privacy concerns remain a critical barrier to real-world deployment. We present a privacy firewall pipeline based on a U-Net encoder-decoder, trained entirely on synthetic data, that removes speech from ambient audio while preserving environmental sounds indicative of daily activities. Activity recognition is performed using VGGish transfer learning with an SVM classifier. Evaluated on the ESC-50 and SINS datasets across multiple speech content levels, the proposed model reduced residual speech to 0% VAD-detectable speech (Silero Voice Activity Detection) under all tested conditions, outperforming Facebook Denoiser (6.55% residual), SepFormer (36.34%) and ConvTasNet (47.21%) on ESC-50 at the 100\% speech level. On ESC-50 at 40% speech level, classification performance recovers to 85% precision and 85% recall after speech removal, compared with 81%/75% before removal and an 84%/83% speech-free baseline. Evaluation on real-world participant home recordings collected with the AudioHive app showed 0% VAD-detectable speech after processing while maintaining 76% precision and recall. The pipeline enables privacy-preserving acoustic sensing without sacrificing activity recognition performance, addressing a key obstacle to the adoption of ambient monitoring in elderly care.

Comments42 pages, 8 figures, 4 tables. Submitted to Pervasive and Mobile Computing. Preprint also available on SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7369016

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑