发表机构
Cribl AI Research Lab(Cribl AI研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体自主运行的成本与安全问题,提出OnTrack流监控机制,经SWE-bench评估可提升失败轨迹检测效果并节省约18%计算资源。
AI 中文摘要
智能体被部署在从行程规划、股票交易到IT事件分类等各类应用中。多数情况下,大语言模型(LLM)智能体自主运行,仅配备极少基于规则的安全保障措施,这会因不可逆操作引发成本与安全问题。现有研究要么通过安全保障智能体监控行为,要么事后评估日志:前者会为每一步增加成本与延迟,后者则在运行结束后才给出结论,此时已消耗令牌且可能造成损害。为解决该问题,我们提出OnTrack,一种流监控机制,它将智能体的步骤与依赖关系和已记录的成功运行进行对比,每步仅需约1毫秒即可向用户发出警报或阻止智能体。我们在三种访问权限递减的场景中研究该问题:完全参考访问(含历史运行与工具模式)、中间访问(仅含工具模式)、无先验知识(仅含生成的步骤日志)。OnTrack的监控能力随数据访问权限降低而减弱,覆盖从计划违规检测到识别循环、停滞及重复工具调用等任务。最后,我们使用SWE-bench轨迹评估OnTrack:基于前8步,我们的方法将失败轨迹排在成功轨迹下方的效果优于内容相似度方法,AUROC提升0.057;采用弃权(不执行)策略时,我们可节省约18%原本会消耗在失败运行上的计算资源,其中83%的被中断运行确实会走向失败(6次弃权中有5次正确)。
英文摘要
Agents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irreversible actions. Recent works resolve this either by using a safeguard agent to monitor behavior or evaluating logs post-hoc. The first adds cost and latency to every step; the second delivers its verdict after the run, when tokens are burned and damage is done. To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step. We study this problem in three regimes of decreasing access: full reference access (historical runs and tool schemas), intermediate access (only tool schemas), and no prior knowledge (only step logs as generated). Expectation of OnTrack's monitoring capabilities reduces as data access drops, ranging from plan violation detection to identifying loops, stalls, and repeated tool calls. Finally, we evaluate OnTrack using SWE-bench trajectories. Based on the first 8 steps, our method ranks failing trajectories below succeeding ones better than content similarity approaches (+0.057 AUROC). With an abort policy, we save about 18% of compute that would be burned on failing runs, where 83% of interrupted runs were actually heading to failure (5 out of 6 aborts were correct).