arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28490cs.CRcs.AI

基于大语言模型(LLM)的软件与系统安全智能体:方法、应用与评估

LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment

Jingjing Nie, Jiawei Guo, Krishna Meda, Haipeng Cai

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过系统综述2023-2026年的同行评审文献,梳理基于LLM的软件与系统安全智能体的技术方法、应用及评估,指出当前智能体权限未受限、行为不可审计的局限,并展望未来研究方向。

中文摘要 AI 辅助

软件与系统安全工作流通常是程序性的:分析师检查异构工件、形成假设、调用工具、解释输出并修订计划。能够在多步骤工作流中进行规划、使用工具、保持状态并修订行动的基于大语言模型(LLM)的智能体,正被迅速用于自动化这类工作。鉴于将安全决策委托给自主系统的后果,理解此类智能体的构建、使用与评估方式至关重要。然而迄今为止,该领域仍缺乏对已开展工作及当前进展的系统性理解:“智能体”一词的应用不一致,应用场景的风险差异显著,评估协议通常不具可比性。为获得该领域全面且连贯的视角,从而为相关未来研究提供信息,本文对2023年至2026年该领域出现的同行评审文献进行了系统性综述,涵盖三个方面:(1)技术方法,包括智能体架构、感知、记忆、推理与规划、动作空间、编排及自我改进;(2)应用,针对所服务的安全任务;(3)评估,包括所考虑的数据集、结果与轨迹指标、安全措施及基准。我们的综合分析显示,该领域已构建出能够执行动作的智能体,但尚未构建出权限受限或行为可审计的智能体。除了知识系统化,我们还将见解扩展到当前方法、应用及评估设计的局限性与面临的挑战,这为潜在有前景的未来研究方向提供了思路。

英文摘要

Software and systems security workflows are typically procedural: analysts inspect heterogeneous artifacts, form hypotheses, invoke tools, interpret outputs, and revise plans. Large language model (LLM)-based agents, which can plan, use tools, retain state, and revise actions across multi-step workflows, are being rapidly adopted to automate this work. Given the consequences of delegating security decisions to autonomous systems, understanding how such agents are built, used, and assessed is crucial. Yet to this date, there remains a lack of systematic understanding of what has been done and how far we are in this field: the term "agent" is applied inconsistently, applications differ sharply in risk, and assessment protocols are often incomparable. To gain a comprehensive and coherent view of this area hence inform relevant future research, this paper provides a systematic literature review of the (1) technical approaches, including agent architecture, perception, memory, reasoning and planning, action space, orchestration, and self-improvement, (2) applications, with respect to the security tasks served, and (3) assessment, including the datasets, outcome and trajectory metrics, safety measures, and baselines considered, over the peer-reviewed literature spanning the emergence of this area (2023--2026). Our synthesis reveals a field that has built agents able to act but not yet agents whose authority is bounded or whose behavior is auditable. In addition to knowledge systematization, we also extend our insights into the limitations of and challenges faced by current approach, application, and assessment designs, which shed light on potentially promising future research directions.

↑