arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AirKey:用于零训练鲁棒PIN推断的多模态声学辅助WiFi感知

AirKey: Multimodal Acoustic-Assisted WiFi Sensing for Zero-Training Robust PIN Inference

BaiChuan Wu, Bin Liu, Xiang Zhang, Zhi Liu, Jie Zhang, Chao Liu, Huan Yan, Meng Li, Fusang Zhang

arXiv 2608.03151首次发表:更新:

AI 中文总结

AirKey是一种多模态感知框架,通过利用IEEE 802.11机制和声学信号辅助WiFi感知,实现零训练鲁棒PIN推断,准确率超现有单模态方案4倍,可在6次尝试内恢复设备解锁PIN,凸显智能界面的隐私漏洞。

AI 中文摘要

通过WiFi感知实现非接触式按键推断凸显了严重的隐私威胁,但其在现实场景中的可行性受到两个基础物理和部署瓶颈的阻碍:获取稳定感知流对网络权限的严格要求,以及纯WiFi信号在快速肌肉记忆式打字过程中固有的“波形融合”歧义。为克服这些限制,我们提出AirKey,一种新颖的跨模态感知框架,可实现高度隐蔽的零训练PIN窃听。首先,为绕过网络部署障碍,AirKey利用基本的IEEE 802.11机制,可预测地从未修改的目标设备引出确认(ACK)响应。通过使用低成本微控制器被动收集这些ACK中的信道状态信息(CSI),AirKey完全无需网络关联即可确保连续的空间感知流。关键的是,为解决WiFi波形融合瓶颈,AirKey引入了跨模态互补机制。通过利用轻量级声学信号作为精确的时间锚点,系统可鲁棒地指导重叠CSI轨迹的分割。这种联合时空融合严格将CSI衍生的空间相似性与声学引导的按键间时间相交。广泛的现实评估表明,AirKey的准确率比最先进的单模态零训练方案高出4倍以上,可在6次尝试内成功恢复设备解锁PIN。最终,这项工作揭示了当代智能界面中的一个关键漏洞,强调了无处不在的多模态感知带来的严重隐私影响。

英文摘要

Contactless keystroke inference via WiFi sensing highlights severe privacy threats, yet its real-world feasibility is hindered by two fundamental physical and deployment bottlenecks: the strict requirement for network privileges to acquire stable sensing streams, and the inherent "waveform fusion" ambiguity of pure WiFi signals during rapid, muscle-memory typing. To overcome these limitations, we propose AirKey, a novel cross-modal sensing framework that achieves highly stealthy, zero-training PIN eavesdropping. First, to bypass network deployment barriers, AirKey exploits fundamental IEEE 802.11 mechanisms to predictably elicit Acknowledgment (ACK) responses from unmodified target devices. By passively harvesting Channel State Information (CSI) from these ACKs using a low-cost microcontroller, AirKey secures a continuous spatial sensing stream entirely without network association. Crucially, to resolve the WiFi waveform fusion bottleneck, AirKey introduces a cross-modal complementarity mechanism. By utilizing lightweight acoustic signals as precise temporal anchors, the system robustly guides the segmentation of overlapping CSI trajectories. This joint spatiotemporal fusion strictly intersects CSI-derived spatial similarities with acoustic-guided inter-keystroke timing. Extensive real-world evaluations demonstrate that AirKey achieves over 4x higher accuracy than state-of-the-art unimodal zero-training schemes, successfully recovering device-unlock PINs within 6 attempts. Ultimately, this work exposes a critical vulnerability in contemporary smart interfaces, underscoring the severe privacy implications of ubiquitous multimodal sensing.

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑