EvoSteer:基于参考锚定信用分配的在线自进化图编排
EvoSteer: Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment
另 3 家 · 查看机构详情
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Dalian University of Technology(大连理工大学)
- Fudan University(复旦大学)
- University of Oxford(牛津大学)
- North China Electric Power University(华北电力大学)
- The University of Texas Health Science Center at Houston(德克萨斯大学休斯顿健康科学中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
EvoSteer提出在线自进化图编排范式,通过锚定轨迹平衡解决信用扩散,并引入验证性技能准入,在十二个数据集上显著优于基线。
中文摘要 AI 辅助
近年来,基于LLM的多智能体系统被广泛应用于将使用工具的智能体编排成可执行的通信图。然而,现有的自进化编排仍面临关键挑战,包括事后进化(仅在轨迹结束后才修订团队)、信用扩散(在混杂基线条件下,每个动作获得相同的终端优势)以及技能准入(未校准且从不淘汰)。为解决这些挑战,我们提出EvoSteer,一种在线自进化图编排的新范式——编排器构建一个运行中的团队,并根据执行特征和学习到的价值估计来修复其看似合理但失败的步骤。为支持这一范式,我们引入锚定轨迹平衡(AnchorTB),一种回归风格的流匹配损失,通过将子轨迹与冻结参考进行平衡,为每个编排动作分配一个系数。基于学习到的流,我们进一步提出验证性技能准入,其中候选技能在晋升前先试用,且仅当配对证据在共享的名义测试预算下通过顺序测试时才晋升。此外,AnchorTB将测量的任务级参考奖励统计与前缀相关修正相结合。在十二个数据集上的实验结果表明,EvoSteer在问答、数学推理、代码生成和交互式决策方面显著优于基线。我们的代码可在以下网址获取:https://this.url。
英文摘要
In recent years, LLM-based multi-agent systems have been widely applied to orchestrate tool-using agents into executable communication graphs. However, existing self-evolving orchestration still faces key challenges, including post-hoc evolution that revises the team only after the trajectory ends, credit diffusion that gives every action the same terminal advantage under confounded baselines, and skill admission that is uncalibrated and never retired. To address these challenges, we propose EvoSteer, a new paradigm of Online Self-Evolving Graph Orchestration -- the orchestrator builds a running team and repairs its plausible but failing steps from execution features and a learned value estimate. To support this paradigm, we introduce Anchored Trajectory Balance (AnchorTB), a regression-style flow-matching loss that assigns each orchestration action a coefficient by balancing subtrajectories against a frozen reference. Built on the learned flow, we further propose Validated Skill Admission, in which a candidate skill is tried before promotion and promoted only if paired evidence passes a sequential test under a shared nominal testing budget. Moreover, AnchorTB combines measured task-level reference reward statistics with prefix-dependent corrections. Experimental results on twelve datasets show that EvoSteer significantly outperforms baselines across question answering, mathematical reasoning, code generation, and interactive decision making. Our code is available at https://github.com/beita6969/evosteer.