弥合首小时缺口:评估执法部门网络事件响应中的人工智能可靠性与基准缺陷
Bridging the First-Hour Gap: Evaluating AI Reliability and Benchmarking Deficiencies in Cyber Incident Response for Law Enforcement
- National Institute of Technology Calicut(卡利卡特国立技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文系统调查了执法部门网络事件响应首小时的决策支持架构,指出RAG系统为可行中间方案,但现有基准不足,需新基准关注新手查询鲁棒性与证据保全。
AI中文摘要:
一线执法人员在网络事件发生后的最初一小时内的行动,对决定调查的最终成败起着至关重要的作用。他们犯下的微小错误可能导致不可逆转的重大影响。由于最初一小时发生的微小错误,调查的完整性可能受到损害,对网络犯罪分子的起诉可能受到阻碍。这主要是因为数字证据的易失性,可能导致程序错误和证据流失。本文对旨在协助网络犯罪第一响应者的决策支持架构进行了系统性调查,将其分类为剧本、大型语言模型(LLMs)、检索增强生成(RAG)框架和智能体人工智能系统。该调查批判性地考虑了实际场景中有限的技术熟练度和不一致的法证基础设施的约束。我们的分析表明,基于RAG的系统因其自然语言适应性而成为一种相对可行的中间解决方案。然而,显著的風險因素,如提示敏感性和在法律语境中自信幻觉的可能性,构成了重大挑战。此外,我们回顾了当前的网络安全基准,并证明它们不足以捕捉执法部门在犯罪最初一小时内的特定安全和法律要求。我们最后论证了需要一个专注于新手查询鲁棒性和证据保全的新评估基准,以确保人工智能驱动的指导符合司法程序的强制性要求。
英文摘要:
The actions of frontline law enforcement officers in the initial hour of a cyber incident play a vital role in determining the ultimate success of an investigation. The minor mistakes they commit might result in irreversible critical impacts. The integrity of the investigation can be compromised, and the prosecution of cyber criminals can be hindered due to minor mistakes that happen in the initial hour. These are mainly because of the volatile nature of digital artifacts that might lead to procedural errors and evidence attrition. This paper provides a systematic survey of decision-support architectures designed to assist first responders of a cybercrime, categorizing them into playbooks, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) frameworks, and Agentic AI systems. The survey critically considers the constraints of limited technical proficiency and inconsistent forensic infrastructure in a practical scenario. Our analysis identifies RAG-based systems as a relatively viable intermediate solution due to their natural language adaptability. However, significant risk factors like prompt sensitivity and the potential for confident hallucinations in legal contexts pose a major challenge. Furthermore, we review current benchmarks in cybersecurity and demonstrate that they are not sufficient to capture the specific safety and legal requirements of law enforcement, focusing on the initial hour of the cybercrime. We conclude by arguing for the necessity of a new evaluation benchmark focused on naive query robustness and evidence preservation, so as to ensure that AI-driven guidance aligns with the mandatory demands of judicial proceedings.