arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TRACE-ROUTER:用于智能AI的任务一致且自适应的在线路由

ORACLE: Agentic AI Orchestrator Routing Via Adaptive Verifier Calibration Feedback

Ritik Raj, Souvik Kundu, Dheemanth Joshi, Tushar Krishna

arXiv 2607.22465首次发表:更新:

发表机构

Georgia Institute of Technology; Intel; Texas A&M University(佐治亚理工学院; 英特尔公司; 德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对企业AI中现有路由决策与智能应用任务级结果不匹配问题,提出TRACE-Router任务级路由框架,利用上下文博弈、终端奖励更新策略,在多基准测试中改善准确性-延迟权衡,取得较好效果。

AI 中文摘要

选择具有不同成本-质量权衡的大语言模型进行路由已成为企业AI的基本部署特征。现有路由器主要为每个大语言模型调用独立做出路由决策。然而,智能应用作为长期工作流执行,其质量仅由延迟的任务级结果决定。这种不匹配使逐调用路由器无法将反馈正确归因于单个路由决策。为缓解此问题,我们提出TRACE-Router,这是一个任务级路由框架,使路由与监督单元对齐。TRACE-Router在接纳时使用上下文博弈为每个任务分配一个模型,将所有后续大语言模型调用固定到所选后端,并使用任务的终端奖励更新其策略,综合考虑准确性和延迟。通过利用延迟任务反馈,TRACE-Router学习适应工作负载的路由策略,同时避免显式任务复杂性估计。在三个智能基准测试中,TRACE-Router持续改善准确性-延迟权衡,实现非支配帕累托前沿点。在tau2-Bench上,它比单个模型之间的延迟匹配插值性能高7-8个准确性点,在Terminal-Bench上,它比最强的单模型基线准确性高7.1个点,延迟低36%。

英文摘要

Modern enterprise agent deployments consist of a heterogeneous pool of large language models (LLMs) having diverse capabilities and cost. Existing model routing strategies optimize the quality-cost trade-off, while providing request-level static decisions. More recent solutions address agentic routing as a task-level selection with a serial verifier based router feedback loop. However, their fixed verifier suitable for homogeneous workloads may not generalize to heterogeneous batches of agentic tasks (example: coding, general conversational). Additionally, due to the verifier placement in the critical path of the loop, serving quality may be affected during multiple concurrent requests routing. To mitigate these issues, we present ORACLE. It is a concurrency-aware online routing mechanism that aligns adaptive routing with adaptive verification for feedback. ORACLE acts as a training-free drop-in 'feedback loop' on top of any model-selection policy to first classify the task type and then dynamically assigns a task-appropriate verifier. Further, we develop a delayed feedback strategy for concurrent requests that largely removes verifier latency from the dispatch critical path. We then present a post-routing dispatch scheduler, namely DISC. DISC reserves each task's peak KV footprint at admission and dispatches to an alternate backend when the reward gain from reduced wait exceeds the reward loss from lower accuracy. Extensive evaluation on SWE-bench, tau2-bench, and Terminal-Bench 2.0 shows that ORACLE improves the accuracy-cost frontier by up to 7 percentage points over state-of-the-art routing baselines, while ORACLE with DISC improves program throughput by up to 1.8x. To facilitate reproducibility and support future development, we open-source the code-base of ORACLE at https://github.com/ORACLE-org/ORACLE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑