发表机构
Stony Brook University; Wuhan University; Shanghai Jiao Tong University(石溪大学; 武汉大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体LLM系统的token通胀问题,提出四阶段路由框架InflationAgent,通过测量通胀、引入CBE信号、最大化SER并采用新鲜升级策略,在GSM8K上优于FrugalGPT。
AI 中文摘要
当语言模型首次尝试无法回答查询时,智能体系统会重试,每次重试都会消耗额外的token。这种重试开销造成了模型的每token价格所暗示的成本与实际完整工作流成本之间的差距,我们将此差距称为“token通胀”,并将其定义为实际工作流成本与单次调用成本的比值。像FrugalGPT这样的系统基于单次调用成本进行路由,在困难任务上可能低估实际成本超过2倍。我们通过InflationAgent解决这一问题,这是一个四阶段路由程序:(1)跨模型层级和任务类型系统地测量token通胀,发现7B模型在多跳问答任务上的通胀高达4.25倍;(2)引入思维链分支熵(CoT Branching Entropy,CBE),这是一种完全从局部推理计算的预执行难度信号,其预测高通胀的AUROC为0.887;(3)通过最大化语义交换率(Semantic Exchange Rate,SER)来选择模型,SER将预期准确率除以预测的实际成本,同时采用新鲜升级策略,在路由到更强模型前丢弃失败的推理链。在固定预算下的GSM8K数据集上,InflationAgent达到94.7%的准确率,而FrugalGPT为91.0%,同时使用的token减少了31%;我们还表明,将失败的推理链转发给GPT-4o会使其准确率降低多达34.8个百分点,验证了新鲜升级设计的有效性。
英文摘要
When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what a model's per-token price implies and what a full workflow actually costs. We call this gap \emph{token inflation} and define it as the ratio of true workflow cost to single-call cost. Systems like FrugalGPT route based on the latter, which can underestimate real cost by more than $2\times$ on difficult tasks. We address this with InflationAgent, a four-stage router that (1) measures token inflation systematically across model tiers and task types, finding inflation as high as $4.25\times$ for a 7B model on multi-hop question answering; (2) introduces CoT Branching Entropy (CBE), a pre-execution difficulty signal computed entirely from local inference, which predicts high inflation with AUROC 0.887; and (3) selects models by maximizing a Semantic Exchange Rate (SER) that divides expected accuracy by predicted true cost, with a fresh-escalation policy that discards failed chains before routing to a stronger model. On GSM8K under a fixed budget, InflationAgent achieves 94.7\% accuracy versus 91.0\% for FrugalGPT while using 31\% fewer tokens, and we show that forwarding a failed reasoning chain to GPT-4o reduces its accuracy by up to 34.8 percentage points, validating the fresh-escalation design.