arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01772cs.AI

FRAMES:面向策略管控型企业工作流中智能体的受保护双目标技能演化

ESCROW: Guarded and Dual-Objective Continual Maintenance for Agents in Policy-Governed Enterprise Workflows

Ruoqi Shu, Chen Dan, Xuhui Wang, Tianhua Xu, Mengxi Luo, Yanming Mai, Bo Wan

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体在受策略约束的企业工作流中改进难的问题,提出FRAMES闭环框架,可从现有资产冷启动技能并通过多机制演化,在保持可审计性的同时实现最优准确性-成本权衡,且结果可在tau-bench复现。

中文摘要 AI 辅助

大型语言模型(LLM)智能体越来越多地运行受策略约束的企业工作流,如文档审核,在此场景中它们必须一致地应用规则、为每个值提供依据并保持可审计性。改进这些智能体颇具挑战:运营反馈稀疏且无标签,对某一规则的修改可能导致不相关案例的性能退化,且必须在不增加推理成本或丧失可审计性的前提下提升准确性。我们提出FRAMES,这是一种闭环框架,可从现有资产冷启动可部署技能,随后通过基于共识的变异、针对准确性与成本的帕累托选择以及反退化保障对技能进行演化,同时全程保持可审计性。在我们的内部生产系统中部署FRAMES后,其在基准方法中实现了最优的准确性-成本权衡,且该结果在tau-bench上得到复现。

英文摘要

LLM agents increasingly run policy-bound enterprise workflows, where they must apply rules consistently and stay auditable. Deploying such an agent is the start of its long-term maintenance cycle: it must adapt to a stream of operational signals, yet reliably turning these sparse, unlabeled signals into reusable skill revisions is hard, and a careless update can trade one task category's accuracy for the overall gain, revive a resolved failure, or land at an undeployable cost. We present ESCROW, a post-deployment maintenance framework that updates an agent's external, reviewable skills under a Strict Update Boundary: the LLM proposes candidate revisions, but only an empirically evaluated version is deployed. It combines distributed diagnosis with consensus, a per-category non-regression guard, cross-cycle anti-regression, and accuracy--cost Pareto search, emitting a versioned, auditable diff per change. In real production on our internal financial document-auditing system, it attains the strongest evaluated accuracy--cost trade-off among baselines, with a transfer probe on public $τ$-bench.

发表机构

  • BMO Financial Group(加拿大蒙特利尔银行金融集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑