arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09087cs.CR

PrivEscalate:测量与增强LLM自动化Linux权限提升的威胁

PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation

Yixuan Liu, Zilong Zhen, Yin Wu, Yi Li

首次发表
浏览论文内容

中文总结 AI 辅助

提出PrivEscalate基准(531个Docker化场景)测量LLM智能体在Linux权限提升中的能力,发现模型能力异质性、环境敏感性和架构影响,并开发PrivEscAgent增强智能体以提升成功率。

中文摘要 AI 辅助

随着大语言模型(LLM)智能体在网络杀伤链中越来越多地自动化进攻性操作,其在复杂本地后渗透任务中的效能仍未得到充分量化。其中,Linux权限提升是从初始访问到完全系统入侵之间的关键步骤。然而,现有针对该任务的评估受限于小样本量(少于15个场景),缺乏在可执行验证下比较模型能力的规模。为解决这一问题,我们提出了PrivEscalate,一个大规模Linux权限提升基准,包含531个Docker化场景,涵盖14个子类别。我们额外衍生出329个参数化变体,以测量对环境干扰物的敏感性。对六种LLM在三种智能体架构下的评估揭示:(i)模型能力在不同漏洞类别间存在异质性,没有单一模型在高流行类别中占主导地位,这促使需要多维风险评估;(ii)LLM的成功对环境扰动敏感,因此配置轮换可以破坏某些利用尝试,但并不能消除所测量的风险;(iii)智能体架构能实质性改变成功率并重新排列模型排名,尽管影响幅度取决于模型。利用这些见解,我们开发了PrivEscAgent,一个领域专用包装器,通过确定性枚举、类别匹配和步骤规划来增强通用ReAct智能体。PrivEscAgent在无需修改底层LLM的情况下,超越了先前的Linux权限提升智能体基线。我们开源PrivEscalate作为Docker化测量工具,支持LLM智能体评估、防御工具验证和红队训练。

英文摘要

As Large Language Model (LLM) agents increasingly automate offensive operations across the cyber kill chain, their efficacy in complex local post-exploitation tasks remains inadequately quantified. Among these, Linux privilege escalation is a key step between initial access and full system compromise. However, existing evaluations for this task are limited by small sample sizes (fewer than 15 scenarios), lacking the scale to compare model capabilities under executable verification. To address this, we present PrivEscalate, a large-scale benchmark for Linux privilege escalation, comprising 531 Dockerized scenarios spanning 14 sub-categories. We additionally derive 329 parameterized variants to measure sensitivity to environmental distractors. Evaluating six LLMs across three agent architectures reveals: (i) model capability is heterogeneous across vulnerability classes, with no single model dominating across the high-prevalence classes, motivating multi-dimensional risk assessments; (ii) LLM successes are sensitive to environmental perturbation, so configuration rotation can disrupt some exploit attempts but does not eliminate the measured risk; and (iii) agent architectures can materially change success rates and reorder model rankings, though the magnitude is model-dependent. Leveraging these insights, we develop PrivEscAgent, a domain-specialized wrapper that augments a generic ReAct agent with deterministic enumeration, category matching, and step planning. PrivEscAgent improves over prior Linux privilege-escalation agent baselines without underlying LLM modifications. We release PrivEscalate as an open-source, Dockerized measurement instrument supporting LLM agent evaluation, defensive tool validation, and red-team training.

发表机构

  • Nanyang Technological University(南洋理工大学)
  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑