arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38954cs.CRcs.AI

APTInvestBench:评估不同遥测条件下的自主APT调查

APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry

Yu Wang, Shuhao Li, Tao Yin, Ziyang Li, Xueying Zhao, Peishuai Sun, Jiang Xie

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体在不同遥测条件下调查APT的鲁棒性问题,提出APTInvestBench基准,含370个案例和1640万条日志,揭示覆盖率掩盖下的证据支持不稳定,助力开发更可靠的防御智能体。

中文摘要 AI 辅助

大语言模型(LLM)智能体可通过将微弱线索转化为入侵范围界定与响应的证据,帮助安全运营中心(SOC)调查高级持续性威胁(APT)。然而,在一种遥测设置下的成功并不能证明其对日志收集、保留或采样变化的鲁棒性。我们引入APTInvestBench,一个用于评估自主APT调查中跨遥测鲁棒性的基准。它包含370个案例,涵盖七种受SOC启发的条件,源自56个基于报告的攻击重建,包含1640万条日志记录。智能体调查未经验证的线索并提交带有记录级引用的报告。固定的行动级支持要求追踪可用日志、查询返回和正式引用中的充分证据,将遥测限制与获取和报告缺口区分开来。在十一个LLM中,智能体平均为44.3%的可恢复攻击行动获取了充分证据,而正式引用仅支持25.0%。更重要的是,总体覆盖率可能掩盖显著的不稳定性:从完整遥测到仅端点遥测,覆盖率仅下降1.6个百分点,但35.5%先前被覆盖的行动失去了充分的引用支持,尽管这些行动仍可恢复。在四个框架中,即使注册的支持记录保持不变,此类损失仍然存在。APTInvestBench提供了可重用的调查环境和诊断性评估,以识别这些缺口并开发更可靠的防御性智能体。

英文摘要

Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response. Yet success under one telemetry setting does not establish robustness to changes in log collection, retention, or sampling. We introduce APTInvestBench, a benchmark for evaluating cross-telemetry robustness in autonomous APT investigation. It comprises 370 cases across seven SOC-inspired conditions, derived from 56 report-informed attack reconstructions with 16.4 million log records. Agents investigate unverified leads and submit reports with record-level citations. Fixed action-level support requirements track sufficient evidence across available logs, query returns, and formal citations, separating telemetry limitations from acquisition and reporting gaps. Across eleven LLMs, agents acquire sufficient evidence for 44.3% of recoverable attack actions on average, while formal citations support only 25.0%. More importantly, aggregate coverage can conceal substantial instability: from Full to endpoint-only telemetry, coverage declines by only 1.6 percentage points, yet 35.5% of previously covered actions lose sufficient citation support despite remaining recoverable. Across four frameworks, such losses persist even when registered supporting records remain unchanged. APTInvestBench provides reusable investigation environments and diagnostic evaluation for identifying these gaps and developing more reliable defensive agents.

发表机构

  • Zhongguancun Laboratory(中关村实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑