arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于智能体轨迹的自动机:失败与下一步预测

Automata from Agent Traces: Failure and Next-Step Prediction

Seonglae Cho, Franklin Cardenoso Fernandez, Umar Mohammed, Zekun Wu, Kleyton Da Costa, Ilham Wicaksono, Adriano Koshiyama

arXiv 2608.23670首次发表:更新:

发表机构

Holistic AI; PUC-Rio; University College London(整体人工智能公司; 里约热内卢天主教大学; 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出将LLM智能体的轨迹语料库压缩为有限状态机(FSM),该模型在12个公开数据集上表现紧凑、准确且构建快速,可同时提升下一步预测与失败预测性能,为安全监控提供模型无关的结构基元。

AI 中文摘要

基于大语言模型(LLM)的智能体可执行多步骤任务,但其行为结构仍不透明:冗长的非结构化轨迹难以满足部署所需的安全审计与运行时监控要求。现有方法仅针对单条轨迹或仅考虑成功案例,因此遗漏了关联下一步预测与失败预测的跨运行拓扑结构。为恢复这种共享结构,我们将整个轨迹语料库压缩为单个紧凑的有限状态机(FSM),作为LLM智能体原本不可预测行为的结构基础。在12个公开数据集上,这些FSM十分紧凑(7至43个状态),在拆分数据集上的重播保留度≥0.997,且各拆分间拓扑结构近乎一致,构建仅需毫秒级时间。该结构基础可同时解决两类预测目标:对于下一步预测,FSM状态上下文在所有与真实值匹配的数据集上的表现均优于Agent Workflow Memory;对于失败预测,每个状态的行为特征在保留数据集上的AUROC最高达0.94,且在线监控器可在部分轨迹阶段将失败运行排在成功运行之上,在任务完成前触发早停。因此,行为拓扑似乎更多由部署框架而非LLM塑造,为安全审计与运行时监控提供了一种模型无关的结构基元。

英文摘要

LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. To recover that shared structure, we collapse an entire trace corpus into a single, compact finite-state machine (FSM) that serves as a structural substrate for the otherwise unpredictable behavior of LLM agents. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness with near-identical topology across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context outperforms Agent Workflow Memory on every ground-truth-matched dataset. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion. Behavioral topology thus appears shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑