arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将按键声音转换为文本:对键盘的自监督声学窃听攻击

Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards

Atsunori Okada, Akira Ito, Rei Ueno, Yuichi Hayashi, Naofumi Homma

arXiv 2607.22094首次发表:更新:

发表机构

Tohoku University; Kyoto University; Nara Institute of Science and Technology(东北大学; 京都大学; 奈良先端科学技术大学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出自监督声学窃听攻击,结合无监督声学聚类、Transformer语言模型推理和迭代自训练,仅通过按键声音重建文本。在多种场景测试中,少量按键就能实现高准确率,揭示了现实的隐私风险。

AI 中文摘要

我们提出了一种自监督声学窃听攻击,仅从按键声音中重建输入的文本,无需目标设备的标记数据。该攻击能在现实世界的两种场景(物理空间和在线会议)中进行隐秘窃听。我们的方法将无监督声学聚类与基于Transformer的语言模型推理和迭代自训练相结合,在高度不确定的声学到字符映射下实现稳定的字符推理。实验表明,在近距离录音设置下,仅用100 - 150次观察到的按键,该方法就能达到超过99%的重建准确率,在低数据量情况下显著优于先前的无监督基线。我们还在多个笔记本电脑平台和现实采集渠道中评估了其鲁棒性,在不同场景下,约150 - 250次观察按键就能达到高重建准确率(常超过90%)。这些结果表明,在仅音频设置下,从按键声音中准确重建文本在实践中是可行的,凸显了一个现实且此前被低估的隐私风险。

英文摘要

We present a self-supervised acoustic eavesdropping attack that reconstructs typed text solely from keystroke sounds, without requiring labeled data for the target device. The proposed attack enables stealthy eavesdropping in two real-world scenarios-physical spaces (public and semi-public) and online meetings. Our method combines unsupervised acoustic clustering with Transformer-based language model inference and iterative self-training, enabling stable character inference under highly uncertain acoustic-to-character mappings. We demonstrate that the proposed method achieves over 99% reconstruction accuracy with only 100-150 observed keystrokes under a close-proximity recording setup using a smartphone placed near the target device, significantly outperforming prior unsupervised baselines in low-data regimes. We further evaluate robustness across multiple laptop platforms and in realistic acquisition channels, including distance recording from approximately 3 meters away on the same desk, through-the-wall eavesdropping with a contact microphone, and background keyboard noise in online conferencing systems. Across these scenarios, the proposed method achieves high reconstruction accuracy (often exceeding 90%) with approximately 150-250 observed keystrokes. These results indicate that accurate text reconstruction from keystroke sounds is feasible in practice under an audio-only setting, even with limited observed keystrokes and without requiring device-specific labeled data, highlighting a realistic and previously underestimated privacy risk.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑