arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GUIAuditor:通过移动设备上的动作引导GUI溯源实现事后儿童安全取证

GUIAuditor: Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices

Junlin Liu, Yifeng Cai, Shuai Wang, Zhineng Zhong, Shaofei Li, Jiacheng Liu, Yuanchun Li, Ziqi Zhang, Xiangqun Chen, Ding Li, Yao Guo

arXiv 2609.28205首次发表:更新:

发表机构

Peking University; Beijing Tongming Lake Information Technology Application Innovation Center; Tsinghua University; University of Illinois Urbana-Champaign(北京大学; 北京通明湖信息技术应用创新中心; 清华大学; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GUIAuditor通过多模态大语言模型和证据蒸馏流水线,在移动设备上创建可查询的GUI溯源记录,实现事后儿童安全取证,显著减少数据分析量并保持高准确率。

AI 中文摘要

智能设备的普及使儿童面临网络风险,如诱骗和金融诈骗,这些风险深嵌于合法应用之中。当前方法依赖自动化预防和检测,这一范式因其固有的易错性而受到根本限制。无论是基于规则还是基于人工智能的方法,都不可避免地产生误报和漏报,无法提供可靠的保护。在本文中,我们主张一种互补的、人在环路的、事后取证范式。我们提出了GUIAuditor,这是首个旨在通过创建GUI溯源(GUI Provenance)来实现这一愿景的系统:一种可查询的、语义化的儿童交互序列记录。为了生成此记录,GUIAuditor利用多模态大语言模型(MLLM)将GUI事件的时间序列转换为人类可理解的叙述。为了在移动设备上实现实用性,一种新颖的证据蒸馏流水线相比行业标准采用的周期性采样方法,将需要分析的数据减少了超过89.2%,而对准确性的影响可忽略不计。在一个包含295个交互片段的新数据集上,GUIAuditor在记录重要事件方面达到了95.23%的宏F1分数,并且关键的是,其两阶段取证查询引擎在超过90.20%的自然语言问题中成功将正确证据作为首要结果检索出来。在三部现代智能手机上的端到端评估表明,包括设备上MLLM推理在内的完整流水线增加了2.1瓦的功耗和每次事件7.4秒的延迟,峰值内存占用约为3.1GB。这些结果表明,事后GUI取证可以在现代移动设备上运行,并为监护人主导的安全审查提供有用的上下文。

英文摘要

The proliferation of smart devices exposes children to online risks like grooming and financial scams that are deeply embedded within legitimate applications. Current approaches rely on automated prevention and detection, a paradigm that is fundamentally limited by its inherent fallibility. Whether rule-based or AI-driven, they inevitably produce false positives and negatives, failing to provide reliable protection. In this paper, we argue for a complementary, human-in-the-loop, post-hoc forensic paradigm. We present GUIAuditor, the first system designed to realize this vision by creating GUI Provenance: a queryable, semantic record of a child's interaction sequence. To generate this, GUIAuditor leverages a Multimodal Large Language Model (MLLM) to translate the temporal sequence of GUI events into a human-understandable narrative. To make this practical on mobile devices, a novel evidence distillation pipeline reduces the data requiring analysis by over 89.2% compared to periodic sampling approaches adopted by industry standards, with negligible impact on accuracy. On a new dataset of 295 interaction clips, GUIAuditor achieves a 95.23% Macro-F1 Score in logging significant events and, crucially, its two-stage forensic query engine successfully retrieves the correct evidence as the top result for over 90.20% of natural language questions. An end-to-end evaluation on three modern smartphones shows that the full pipeline, including on-device MLLM inference, adds 2.1W of power draw and 7.4s of per-event latency, with a peak memory footprint of ${\sim}$3.1GB. These results show that post-hoc GUI forensics can run on modern mobile devices and provide useful context for guardian-led safety review.

CommentsAccepted by ACM IMWUT/Ubicomp 2026

DOI:10.1145/3831978

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑