arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TRIAGE:用于高效执行的三级路由与智能体引导框架

TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution

Ruocan Wei

arXiv 2609.01428首次发表:更新:

发表机构

China Telecom Cloud(中国电信云)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对ReAct范式LLM智能体的效率问题,提出TRIAGE三级路由框架,通过复用历史轨迹实现三级查询处理,在安全监控和跨域实验中显著减少token消耗,形成效率提升的正反馈循环。

AI 中文摘要

基于ReAct范式的大语言模型(LLM)智能体在工具使用和任务执行方面展现出卓越能力。然而ReAct存在根本性效率问题:每次查询都会从头触发完整推理循环,相似查询会重复相同步骤而不利用历史经验。我们提出TRIAGE,一种通过复用历史执行轨迹减少token消耗的三级路由框架,其核心创新是TaaS(Trajectory-as-a-Skill,将历史执行轨迹抽象为可复用技能,实现“经验即服务”)。TRIAGE将查询分为三级:(1)直接复用——相同查询,消耗0个token;(2)技能替换——相似查询,通过确定性参数替换消耗0个token;(3)完整ReAct——新颖查询,自动存储以供未来复用。在1007个安全监控查询的大规模实验中,TRIAGE实现62.3%的token节省,其中56.0%的查询属于Level 2、5.5%属于Level 1,均零成本执行;在ToolBench(15个领域、345个查询)的跨域验证中,实现76.3%的token减少,证实语义路由的通用性。在线学习实验显示出从冷启动到成熟的演化:前100个查询内Level 2命中率从0%升至57%,平均token成本从198降至74.7。我们还提出自动技能提取机制,将高频轨迹模式提炼为确定性技能,形成“使用越多、效率越高”的正反馈循环。

英文摘要

Large Language Model (LLM) agents based on the ReAct paradigm have demonstrated remarkable capabilities in tool use and task execution. However, ReAct suffers from a fundamental efficiency problem: every query triggers a complete reasoning loop from scratch, and similar queries repeat identical steps without leveraging historical experience. We propose TRIAGE,a three-level routing framework that reduces token consumption by reusing historical execution trajectories. Its core innovation is TaaS (Trajectory-as-a-Skill), which abstracts historical execution trajectories into reusable skills, realizing 'experience as a service'. TRIAGE classifies queries into three levels: (1) Direct Reuse-identical queries, 0 tokens; (2) Skill Substitution-similar queries, 0 tokens via deterministic parameter substitution; (3) Full ReAct-novel queries, automatically stored for future reuse. In large-scale experiments on 1,007 security monitoring queries, TRIAGE achieves 62.3% token savings, with 56.0% of queries at Level 2 and 5.5% at Level 1, both executing at zero cost. Cross-domain validation on ToolBench (15 domains, 345 queries) achieves 76.3% token reduction, confirming the generalizability of semantic routing. An online learning experiment demonstrates cold-start-to-mature evolution: the L2 hit rate rises from 0% to 57% within the first 100 queries, and the average token cost drops from 198 to 74.7. We also propose an automatic Skill extraction mechanism that distills high-frequency trajectory patterns into deterministic Skills, creating a positive feedback loop of 'the more you use it, the more efficient it becomes'.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑