面向自主大语言模型智能体中结构化漂移诊断与恢复的基于图的强化学习框架
A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents
AI总结:
本研究针对自主LLM智能体的运行时行为漂移问题,提出基于图的强化学习即插即用恢复框架,经AppWorld基准验证,该框架可利用漂移信息做出正确恢复决策且输出符合要求。
AI中文摘要:
自主大语言模型(LLM)智能体正越来越多地部署在复杂的现实工作流程中,但它们仍易受到运行时行为漂移的影响,这种漂移是指与原始任务的偏离,可能会对外部系统造成不可逆的副作用。现有方法在提示层面处理漂移,但缺乏用于步骤级检测、风险评估和恢复决策的结构化机制。由于执行主要任务的智能体通常是一个庞大且昂贵的模型,无法在每次部署时都进行重新训练,因此本研究的目标是开发一个即插即用的恢复模块。它引入了一个基于图的框架,其中一个小型语言模型通过强化学习进行训练,专门处理主智能体外部的恢复图的每个节点。每个节点都有明确的角色:漂移分类、操作检测、风险评估或最终决策,模型学习生成适合该角色的结构化XML格式推理。训练结合了基于规则的结构奖励和“LLM作为评判者”的语义质量信号,使模型既根据回答方式(模式和长度)又根据回答内容进行评分。在公开的AppWorld基准上进行的实验表明,该方法通常会利用疑似漂移发生的信息,使用小型语言模型做出正确的恢复决策。此外,经过训练的小型语言模型能可靠地遵守规定的输出模式,并根据其分配的节点角色在每个字段中生成语义恰当的内容。
英文摘要:
Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on external systems. Existing approaches address drift at the prompt level but lack structured mechanisms for step-level detection, risk assessment, and recovery decision. Because the main task-executing agent is often a large and expensive model that cannot be re-trained on every deployment, this work targets a plug-and-play recovery module instead. It introduces a graph-based framework in which a single small language model is trained via reinforcement learning to specialize at each node of a recovery graph, external to the main agent. Each node has a precise role\,: drift classification, operation detection, risk evaluation, or final decision and the model learns to produce structured XML-formatted reasoning adapted to that role. Training combines rule-based structural rewards with an LLM-as-judge semantic-quality signal, so that the model is graded both on how it answers (schema and length) and on what it says. Experiments on the public AppWorld benchmark show that the method generally exploits information about the suspected drift onset to issue correct recovery decisions using a small language model. In addition, the trained small language model reliably respects the prescribed output schema and produces semantically appropriate content in each field according to its assigned node role.