FRAMES:面向策略管控型企业工作流中智能体的受保护双目标技能演化
ESCROW: Guarded and Dual-Objective Continual Maintenance for Agents in Policy-Governed Enterprise Workflows
浏览论文内容
中文总结 AI 辅助
针对LLM智能体在受策略约束的企业工作流中改进难的问题,提出FRAMES闭环框架,可从现有资产冷启动技能并通过多机制演化,在保持可审计性的同时实现最优准确性-成本权衡,且结果可在tau-bench复现。
中文摘要 AI 辅助
大型语言模型(LLM)智能体越来越多地运行受策略约束的企业工作流,如文档审核,在此场景中它们必须一致地应用规则、为每个值提供依据并保持可审计性。改进这些智能体颇具挑战:运营反馈稀疏且无标签,对某一规则的修改可能导致不相关案例的性能退化,且必须在不增加推理成本或丧失可审计性的前提下提升准确性。我们提出FRAMES,这是一种闭环框架,可从现有资产冷启动可部署技能,随后通过基于共识的变异、针对准确性与成本的帕累托选择以及反退化保障对技能进行演化,同时全程保持可审计性。在我们的内部生产系统中部署FRAMES后,其在基准方法中实现了最优的准确性-成本权衡,且该结果在tau-bench上得到复现。
英文摘要
LLM agents increasingly run policy-bound enterprise workflows, where they must apply rules consistently and stay auditable. Deploying such an agent is the start of its long-term maintenance cycle: it must adapt to a stream of operational signals, yet reliably turning these sparse, unlabeled signals into reusable skill revisions is hard, and a careless update can trade one task category's accuracy for the overall gain, revive a resolved failure, or land at an undeployable cost. We present ESCROW, a post-deployment maintenance framework that updates an agent's external, reviewable skills under a Strict Update Boundary: the LLM proposes candidate revisions, but only an empirically evaluated version is deployed. It combines distributed diagnosis with consensus, a per-category non-regression guard, cross-cycle anti-regression, and accuracy--cost Pareto search, emitting a versioned, auditable diff per change. In real production on our internal financial document-auditing system, it attains the strongest evaluated accuracy--cost trade-off among baselines, with a transfer probe on public $τ$-bench.
发表机构
- BMO Financial Group(加拿大蒙特利尔银行金融集团)
机构由 AI 辅助整理,请以论文原文为准。