发表机构
Bucknell University(巴克内尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过越南语击键数据和行为操纵威胁模型,评估四种击键建模方法,发现序列模型更优,但检测鲁棒性不足,对抗训练可提升性能。
AI 中文摘要
我们研究了击键动力学在检测大型语言模型(LLM)辅助写作中的鲁棒性。我们引入了一个越南语击键数据集,捕捉了真实的写作模式,包括真实创作、转录和改写。我们还定义了一个基于行为的威胁模型,其中用户故意改变打字模式。为了实施该威胁模型,我们创建了行为操纵的数据变体,旨在规避基于击键的检测。我们在用户独立和上下文独立的设置下,评估了四种击键建模方法:时间和节奏表示,以及使用一维卷积神经网络(1D-CNN)和TypeNet建模的序列表示。结果表明,在大多数情况下,序列模型优于基于特征的方法,并且击键信号编码了关于写作过程的判别性信息。然而,检测并非统一鲁棒:转录被可靠识别,而改写和对抗性操纵的样本在未明确建模时经常被误分类为真实创作。为了解决这个问题,我们结合了使用行为操纵数据的对抗训练,这显著提高了可分离性和鲁棒性。这些结果表明,基于击键的检测关键依赖于对多样化写作行为的暴露,并且在有限条件下的强性能不会在没有针对性建模的情况下泛化到现实或对抗性设置。
英文摘要
We study the robustness of keystroke dynamics for detecting large language model (LLM)-assisted writing. We introduce a Vietnamese keystroke dataset capturing realistic writing modes, including bona fide composition, transcription, and paraphrasing. We also define a behaviorally grounded threat model in which users deliberately alter typing patterns. To implement the threat model, we create behaviorally manipulated variants of the data designed to evade keystroke-based detection. We evaluate four keystroke modeling approaches: temporal and rhythmic representations, and sequential representations modeled with a one-dimensional convolutional neural network (1D-CNN) and TypeNet, under user-independent and context-independent settings. The results show that sequential models outperform feature-based approaches in most cases and that keystroke signals encode discriminative information about the writing process. However, detection is not uniformly robust: transcription is reliably identified, while paraphrasing and adversarially manipulated samples are frequently misclassified as bona fide when not explicitly modeled. To address this, we incorporate adversarial training using behaviorally manipulated data, which substantially improves separability and robustness. These results suggest that keystroke-based detection depends critically on exposure to diverse writing behaviors, and that strong performance under limited conditions does not generalize to realistic or adversarial settings without targeted modeling.
Comments9 pages, 2 figures. Thanh Dong and An Ngo contributted equally. Accepted at the 2026 IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2026)