发表机构
National Research Council; University of Windsor(加拿大国家研究委员会; 温莎大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究诊断多阶段LLM网络智能体在横向移动中的瓶颈,发现验证器不够具体且乐观,瓶颈集中于凭证与横向移动任务,强调评估需涵盖结果、证据、资源与失败适应。
AI 中文摘要
基于多阶段大语言模型(LLM)的网络智能体在完成攻击工作流时,可能仍然脆弱、成本高昂,或依赖于对执行证据的错误解读。仅凭成功率会掩盖效率低下、通过重试进行适应以及成功或失败的识别等问题。我们针对一个自主对抗系统进行了端到端的诊断研究,该系统在企业级横向移动场景中集成了编排器、执行器和验证器三种大语言模型。我们在两个场景和三种模式(专家定义、自我脚手架和完全自主)下评估了六个前沿模型。我们评估了验证器的一致性和证据依据;引入了一个基于子任务条件、成本感知的评分,用于衡量异常令牌使用、重试和运行时间;并使用比较性LLM作为评判者的分析来识别规划缺陷,包括工具错位、计划相似性、过度具体化、探测不足和恢复能力弱。验证器通常具有相关性和证据依据,但往往不够具体且过于乐观。瓶颈集中在凭证和横向移动任务中,随场景复杂性而扩散,并在完全自主模式下变化更大。可靠的评估必须评估结果、证据解读、资源使用以及失败后的适应能力。
英文摘要
Multi-stage LLM-based cyber agents may complete attack workflows while remaining brittle, costly, or reliant on incorrect interpretations of execution evidence. Success rates alone obscure inefficiency, adaptation through retries, and recognition of success or failure. We present an end-to-end diagnostic study of an Autonomous Adversary system with orchestrator, executor, and validator LLMs in enterprise-like lateral-movement scenarios. Six frontier models are evaluated across two scenarios and three modes: expert-defined, self-scaffolded, and fully autonomous. We assess validator consistency and evidence grounding; introduce a subtask-conditioned, cost-aware score for abnormal token use, retries, and runtime; and use comparative LLM-as-a-Judge analysis to identify planning deficiencies, including tool misalignment, plan similarity, over-specification, inadequate probing, and weak recovery. Validators are generally relevant and evidence-grounded but often nonspecific and overly optimistic. Bottlenecks cluster in credential and lateral-movement tasks, spread with scenario complexity, and vary more under full autonomy. Reliable evaluation must assess outcomes, evidence interpretation, resource use, and adaptation after failure.
CommentsThe 19th International Symposium on Foundations & Practice of Security (FPS 2026)