发表机构
TokenRhythm Technologies(TokenRhythm科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大型语言模型智能体中模型选择问题,提出原生执行框架智能体路由范式,依执行框架状态选模型,其决策产生的数据记录形成数据飞轮,经实例化验证,表明该路由不仅控成本,更是智能体训练的数据引擎。
AI 中文摘要
大型语言模型智能体越来越多地通过管理观察、上下文、控制、行动、状态和验证的执行框架来执行。同时,前沿和开放模型在结构上变得更加专业化。现有路由方法大多优化单轮成本-质量权衡,忽略了使智能体不同于聊天完成的执行状态、中间故障和反馈循环。我们提出了原生执行框架智能体路由,这是一种逐步骤路由范式,根据完整的执行框架状态,选择单个最佳拟合模型以实现经济高效的执行,或选择多个互补模型以提高集成式准确性。关键在于每个路由决策自然产生一个结构化数据记录,这些记录形成一个原生执行框架数据飞轮。我们在OpenSquilla中实例化了这个想法,报告研究了在包括DRACO和PinchBench在内的智能体基准测试上的单例和多模型路由,并认为智能体路由不仅是成本控制,更是智能体原生训练的数据引擎。
英文摘要
Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification. At the same time, frontier and open models are becoming structurally specialized: a model that is strong at code editing, long-context recovery, tool use, mathematical reasoning, or low-latency response may not dominate on the other axes. This makes model selection inside an agent a core systems problem rather than a per-query serving trick. Existing routing methods mostly optimize single-turn cost-quality trade-offs and therefore miss the execution state, intermediate failures, and feedback loops that make agents different from chat completion. We propose Harness-Native agentic routing, a step-level routing paradigm that selects either a single best-fit model for cost-effective execution or multiple complementary models for ensemble-style accuracy improvement, conditioned on the full harness state. The key insight is that every routing decision naturally produces a structured data record -- consisting of the query, harness state, model choice or model set, execution trace, outcome, and cost -- whose labels are supplied by the environment rather than by the router itself. These records form a harness-native data flywheel: execution traces train better routers and harness-native models, which improve cost-quality trade-offs and generate more traces under the same budget. We instantiate this idea in OpenSquilla with a four-layer routing stack, an open LightGBM cold-start ranker, and a staged router-model path that turns logged arena records into progressively stronger routing policies. The report studies singleton and multi-model routing on agentic benchmarks including DRACO and PinchBench, and argues that agentic routing is not merely cost control, but a data engine for agent-native training.
CommentsCode: https://github.com/opensquilla/opensquilla