arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EASy:迈向高效的基于大语言模型的智能体系统

E$^3$-Orch: Towards Effective, Efficient, and Extensible Agentic Orchestration with Reinforcement Learning

Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari

arXiv 2608.04588首次发表:更新:

发表机构

Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EASy是一种可训练的智能体框架,通过强化学习优化任务性能与计算效率,采用里程碑-规划-执行工作流,在多类基准上实现了更优的性能-效率权衡。

AI 中文摘要

智能体系统作为一种有前景的范式,通过协调专门的基于大语言模型(LLM)的智能体来解决复杂任务。然而,现有多数系统主要优化任务成功率,却较少考虑执行器能力、计算成本等实际约束下的执行效率。现有基于路由器的方法在对丰富、动态变化的任务上下文、多步骤依赖关系及中间执行反馈进行推理时能力有限,且对未见过的执行器泛化能力较差。我们提出EASy,这是一种可训练的智能体框架,通过强化学习同时优化任务性能与计算效率。EASy为基于LLM的协调器配备异构执行器的能力与成本概况的显式知识,使其能实现超越仅基于性能的路由的上下文感知协调。它还引入里程碑-规划-执行工作流,将复杂任务分解为可管理的里程碑,构建感知依赖关系的执行图,分配合适的执行器,并行化独立步骤,并根据中间结果调整后续决策。为训练协调器,我们开发树状展开流程,探索替代的里程碑分解与执行计划,同时采用多组件奖励,涵盖任务正确性、执行效率与轨迹完整性。在数学推理、具身决策及深度研究基准上开展的大量实验表明,EASy相较于强大的智能体基线,始终能实现更优的性能-效率权衡。

英文摘要

Agentic orchestration enables multiple autonomous agents to solve complex tasks through adaptive decomposition, delegation, and execution. However, existing orchestrators often rely on hand-crafted logic and prompting strategies, limiting adaptation and generalization across tasks and executor configurations. We propose E$^3$-Orch, a reinforcement learning framework for effective, efficient, and extensible agentic orchestration based on a milestone-plan-act workflow. Instead of planning all subtasks upfront, E$^3$-Orch organizes execution around milestones, a mid-level abstraction that scopes planning around meaningful intermediate objectives and allows orchestration decisions to adapt as execution progresses. For each milestone, the orchestrator builds a dependency-aware plan, assigns subtasks to suitable executors, and executes independent subtasks in parallel. We train the orchestration policy from execution feedback, using milestone and plan decisions as units for fine-grained credit assignment. Tree-structured rollouts compare alternative decisions under shared execution histories, while complementary rewards optimize task performance, execution cost, and planning completeness, including an uncertainty-aware performance reward for stochastic downstream outcomes. Across seven benchmarks, E$^3$-Orch achieves the best task performance under multiple executor configurations, improving over the strongest baselines by $0.7$--$3.8$ points and delivering $1.16$--$1.59\times$ higher intelligence efficiency. The learned policy also transfers to unseen executor configurations introduced only at evaluation time, supporting extensible agentic orchestration.

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑