Watch and Crack: 从智能眼镜视频中推断密码
Watch and Crack: Password Inference from Smart-Glasses Video
浏览论文内容
中文总结 AI 辅助
本文提出首个通用视频击键推断流程,利用智能眼镜摄像头以亚像素精度跟踪手指,结合自监督神经网络和泄露密码统计,在数小时至数天内破解长达16字符的人为密码,熵减少高达60比特。
中文摘要 AI 辅助
在公共场所的智能手机上输入密码,使用户面临基于视频的侧信道攻击,攻击者记录输入过程并通过分析手指运动来重建所输入的密码。先前的基于视频的击键推断攻击,针对平板上的自由文本、数字PIN码、图案锁和密码,但都是在不切实际的简化条件下进行的。我们提出了第一个通用的基于视频的击键推断流程,能够从智能手机QWERTY键盘中恢复基于规则的字母数字密码,覆盖所有四种键盘布局,且无需对受害者或其设备做出任何假设。利用流行的Meta Ray-Ban智能眼镜的内置摄像头,我们的流程以亚像素分辨率跟踪设备和打字手指,通过自监督神经网络集成预测击键,动态估计屏幕上的键盘布局,并计算每个键的概率分布。这些从视频中得出的概率与基于21.9亿个泄露密码的先验密码输入分布相结合,以估计所观察密码的排名,从而估计其破解时间。在一项使用符合典型企业密码策略的密码的用户研究中,人为选择的密码长达16个字符可在数小时至数天内被破解,而密码管理器选择的随机密码长达13个字符在现代GPU上不到一小时即可破解,相对于先前的基线,密码熵最多减少60比特。对于人为选择的密码,将视频证据与先验统计相结合,与单独使用任一来源相比,额外实现了约20比特的统计显著减少。该攻击在长达1.8米的距离内成功,适用于不同的攻击者-受害者姿势和视角,并推广到三款流行的智能手机。
英文摘要
Typing passwords on smartphones in public places exposes users to video-based side-channel attacks, in which an adversary records the typing session and reconstructs the entered password by analyzing finger movements. Prior video-based keystroke inference attacks have targeted free-form text on tablets, numeric PINs, pattern locks, and passwords under unrealistically simplified conditions. We present the first general video-based keystroke inference pipeline that recovers rule-based alphanumeric passwords from smartphone QWERTY keyboards, covering all four keyboard layouts and requiring no assumptions about the victim or their device. Using the built-in camera of the popular Meta Ray-Ban smart glasses, our pipeline tracks the device and typing fingertips at sub-pixel resolution, predicts keystrokes via a self-supervised ensemble of neural networks, dynamically estimates the on-screen keyboard layout, and computes per-key probability distributions. These video-derived probabilities are combined with prior password-typing distributions based on 2.19 billion leaked passwords to estimate the rank, and therefore the crack time, of the observed password. In a user study using passwords that adhere to a typical enterprise password policy, human-chosen passwords of up to 16 characters can be cracked in hours to days, and password-manager-chosen random passwords of up to 13 characters in under an hour on a modern GPU, reducing password entropy by up to 60 bits relative to the prior baseline. For human-chosen passwords, combining video evidence with prior statistics yields an additional statistically significant reduction of approximately 20 bits compared to either source alone. The attack succeeds at distances up to 1.8 m, across diverse attacker-victim postures and viewing angles, and generalizes across three popular smartphones.
发表机构
- Tel Aviv University(特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。