发表机构
TU Wien; TU Graz(维也纳技术大学; 格拉茨技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出ATLAS方法,结合轨迹抽象与自动机学习从智能体轨迹推断有限状态模型,以实现对LLM智能体行为的可解释分析、知识迁移及简洁解释。
AI 中文摘要
基于大语言模型(LLM)的智能体正越来越多地被用于软件测试、网络安全评估等复杂任务。尽管这些智能体展现出令人印象深刻的能力,但其行为却难以理解、解释和分析。现有评估主要聚焦于任务成功与否与执行轨迹,对智能体所采用策略的洞察十分有限。本文提出ATLAS(Automata Learning for Agent Trajectory Analysis and Strategy Discovery,即智能体轨迹分析与策略发现的自动机学习),这是一种从智能体轨迹中恢复可解释行为模型的方法。ATLAS将轨迹抽象与自动机学习相结合,以推断能捕捉观测到的智能体-环境交互策略的有限状态模型。这些模型提供人类可理解的洞察,并支持对重复行为、决策点、成功任务完成路径及失败循环的自动化分析。作为概念验证,我们将ATLAS应用于基于LLM的渗透测试智能体生成的轨迹,所得模型揭示了利用易受攻击机器的高级行为策略,而这些策略仅从原始执行轨迹中难以识别。我们讨论了所学行为模型如何支持智能体系统的可解释性、模型引导探索、审计与分析,还展示了从强大的前沿模型到紧凑语言模型的基于符号模型的知识迁移,此外,在包含12台易受攻击机器的渗透测试案例研究中,我们说明模型转换如何生成智能体行为的简洁解释。ATLAS凸显了模型驱动工程的新机遇:将智能体轨迹转换为显式行为模型,从而实现对原本不透明的AI智能体的系统理解与分析。
英文摘要
Large Language Model (LLM)-based agents are increasingly used for complex tasks such as software testing and cybersecurity assessment. While these agents demonstrate impressive capabilities, their behavior is difficult to understand, explain, and analyze. Existing evaluations focus mainly on task success and execution traces, offering limited insight into the strategies employed by the agent. We present ATLAS (Automata Learning for Agent Trajectory Analysis and Strategy Discovery), an approach for recovering interpretable behavioral models from agent trajectories. ATLAS combines trace abstraction with automata learning to infer finite-state models that capture observed agent-environment interaction strategies. These models provide human-interpretable insights and support automated analyses of recurring behaviors, decision points, successful task-completion paths, and failure loops. As a proof of concept, we apply ATLAS to trajectories generated by an LLM-based penetration-testing agent. The resulting models expose high-level behavioral strategies for exploiting vulnerable machines that are difficult to identify from raw execution traces alone. We discuss how learned behavioral models can support explainability, model-guided exploration, auditing, and analysis of agentic systems. We further demonstrate symbolic model-based knowledge transfer from powerful frontier models to compact language models. In addition, we show how model transformations can derive concise explanations of agent behavior in a penetration-testing case study comprising 12 vulnerable machines. ATLAS highlights a new opportunity for model-driven engineering: transforming agent trajectories into explicit behavioral models that enable systematic understanding and analysis of otherwise opaque AI agents.
Comments7 pages, accepted for publication at ACM/IEEE MODELS 2026