arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TROVE:基于轨迹的路径验证与编辑实现自适应智能体技能编排

TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing

Tianxing Wang, Mingming Zhao, Shuai Huang, Huiyang Xu, Chaoyue Niu, Shengzhong Liu, Fan Wu

arXiv 2609.05019首次发表:更新:

AI 中文总结

TROVE是一种自适应智能体技能编排方法,通过离线提炼轨迹生成技能与转换图、在线选择性编辑失效路径,在多类基准上实现了更优的质量-效率权衡。

AI 中文摘要

智能体往往会在观察到决定性运行时结果之前优化、选择或约束执行结构。然而,这种执行前的承诺会造成编排瓶颈:当中间证据使待执行的后续步骤失效时,智能体要么执行过时的步骤,要么进行大范围重规划,这会加剧错误、浪费计算资源并丢弃已取得的进展。为此,我们提出了基于轨迹的验证与编辑的路径编排方法TROVE,它仅修正运行时证据失效的部分。在离线阶段,TROVE将已评估的工作流搜索轨迹提炼为原子技能和复合技能,以及结果条件转换图,保留稳定片段的同时揭示结果依赖的决策。在线阶段,它将规划的路径视为临时方案:在提交一个顶级技能后,控制器会保留有效的后续步骤、插入轨迹支持的局部响应,或仅替换失效的后缀。在代码生成、问答和数学推理基准上,使用不同大语言模型(LLM)骨干的评估显示,TROVE比现有的数据集级优化、查询级架构选择和图约束调度基线实现了更优的质量-效率权衡。当结果改变合适的后续步骤时,质量提升最大;而对于接近饱和的任务,早期终止能带来显著的效率提升。消融实验进一步表明,复合技能捕获了大部分离线收益,插入支持局部修正,后缀替换主要提升效率。这些发现确立了选择性路径编辑是自适应智能体编排的通用原则。

英文摘要

Agents tend to optimize, select, or constrain execution structures before decisive runtime outcomes are observed. However, such pre-execution commitment creates an orchestration bottleneck: when intermediate evidence invalidates the pending continuation, agents must either execute stale steps or replan broadly, compounding errors, wasting computation, and discarding progress. We thus propose Trace-grounded Route Orchestration via Validation and Editing (TROVE), which revises only what runtime evidence invalidates. Offline, TROVE distills evaluated workflow-search traces into atomic and composite skills and an outcome-conditioned transition graph, preserving stable fragments while exposing outcome-dependent decisions. Online, it treats a planned route as provisional: after committing one top-level skill, the controller retains a valid continuation, inserts a trace-supported local response, or replaces only the invalid suffix. Evaluation across code-generation, question-answering, and math reasoning benchmarks with different LLM backbones show that TROVE delivers a stronger quality-efficiency trade-off than existing baselines of dataset-level optimization, query-level architecture selection, and graph-constrained scheduling. Quality gains are largest when outcomes change the appropriate continuation, whereas early termination yields substantial efficiency gains on near-saturated tasks. Ablations further show that composite skills capture most offline benefits, insertion enables local correction, and suffix replacement primarily improves efficiency. These findings establish selective route editing as a general principle for adaptive agent orchestration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑