arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02464cs.AIcs.LGcs.SE

LLM智能体故障的实时检测与修复

Real-Time Detection and Repair of LLM Agent Failures

Sunny Dubey

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出基于LLM智能体步骤遥测的实时故障检测与修复系统,结合单类回声状态网络集成与确定性验证层,可高效检测并修复故障,提升任务成功率且成本极低。

中文摘要 AI 辅助

LLM智能体在运行过程中会出现故障,包括循环、级联工具错误、偏离目标、生成虚假结果或静默吸收损坏内容,而采用第二个LLM对每一步进行评判的标准补救措施,其成本超过了智能体本身。本文仅利用可观测的步骤遥测数据,研究仅通过监控器(每步耗时微秒级,仅在健康运行数据上训练)可实现的故障检测程度。在三个框架、三个本地模型(qwen2.5 7b/3b、llama3.1 8b)以及商用API(gemini-2.5-flash)的2823个已提交智能体运行实例上,采用CUSUM告警的单类回声状态网络集成模型,在5%的误报预算下可检测到0.71的故障,AUROC为0.872。其相对于无记忆基线的优势是故障发生后时间范围的单调函数(<=3步时提升0.09,>=9步时提升0.40),在AFTraj-2K上对自身故障区域的样本外预测表现良好,无需重新训练即可将性能迁移至其他两组语料库(AFTraj-2K为0.745,ATBench为0.779)。监控器存在两个缺陷:每个部署需专属的健康空模型(无法迁移,冷启动时AUROC为0.527,重新校准后为0.885),以及存在残留误报率。本文添加了一个无此缺陷的确定性验证层,该层会根据智能体实际接收的工具结果重新计算运行的总结果,并确认所有必要调用均已执行。对比实验显示,该层可捕获60%的故障(加入覆盖率检查后为96%),在63个健康实例中误报为0,而监控器的误报率为17%;该层无需修改即可迁移至llama3.1:8b(110个实例全部正确,误报为0),在1825个健康实例中误报为0。本文将检测与修复结合:对每个标记的运行进行回滚并重新实时运行,相对于16%的重采样对照组,可恢复45%的故障(p=0.0005),使任务成功率从52%提升至73%,每次运行仅增加约1次模型调用。该系统每步耗时约200微秒,比评判调用快三个数量级,代码、轨迹和结果已公开。

英文摘要

LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself. We ask how much detection is achievable from observable step telemetry alone, using monitors costing microseconds per step and trained only on healthy runs. On 2,823 committed agent episodes across three frameworks, three local models (qwen2.5 7b/3b, llama3.1 8b) and a commercial API (gemini-2.5-flash), a one-class echo-state-network ensemble with CUSUM alarms detects 0.71 of failures at a 5% false-alarm budget (AUROC 0.872). Its advantage over a memoryless baseline is a monotone function of post-onset horizon (+0.09 at <=3 steps, +0.40 at >=9), predicting its own failure region out of sample on AFTraj-2K. Ranking transfers with no retraining to two corpora from other groups (AFTraj-2K 0.745, ATBench 0.779). Monitors carry two burdens: a per-deployment healthy null (they do not transfer -- AUROC 0.527 cold against 0.885 recalibrated) and a residual false-alarm rate. We add a layer carrying neither: deterministic verification, which recomputes a run's stated total from the tool results it actually received and confirms every required call was made. Head-to-head it catches 60% of failures (96% with the coverage check) at 0 of 63 false positives against the monitor's 54% at 17%, transfers unchanged to llama3.1:8b (110 of 110 at 0 of 10), and trips on 0 of 1825 healthy episodes. Detection is then closed into repair: each flagged run is rolled back and re-run live, recovering 45% of failures against a 16% resampling control (p=0.0005) and lifting task success from 52% to 73% for about one extra model call per run. The system runs at ~200 microseconds per step, three orders of magnitude below a judge call. Code, traces and results are released.

补充信息

↑