发表机构
Sun Yat-Sen University; Rensselaer Polytechnic Institute; Southeast University; Hong Kong University of Science and Technology(中山大学; 伦斯勒理工学院; 东南大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TokenCast通过可组合成本表示预测LLM智能体执行中的令牌消耗,动态更新预测,显著降低误差并节省令牌。
AI 中文摘要
当大型语言模型(LLM)智能体执行同一任务时,不同运行间的令牌消耗可能相差一个数量级以上。智能体根据工具反馈和中间结果选择下一步行动,而不断增长的上下文会稳定地增加每次后续调用的输入大小。因此,任务的总消耗在执行前难以预测,且预测必须随着运行过程不断修正。在本文中,我们提出TokenCast,它为每个执行片段学习一个可组合的成本表示,记录其自身消耗以及其引入的上下文增长。组合相邻片段可得到累积估计,该估计捕捉到早期片段的上下文被后续每次调用重新读取时产生的额外输入成本。随着执行展开,新观察到的证据会刷新预测,无需额外的LLM调用,在SWE-bench Verified上平均每次运行的累积预测时间为32.8毫秒。在4个任务套件和6个智能体模型上,TokenCast相对于最强比较器的平均绝对误差降低平均为14.5%(在96个评估组合上)。在离线预算控制重放中,TokenCast在匹配的轨迹完成度下平均比固定预算策略少使用21.3%的令牌。代码可在该https URL获取。
英文摘要
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.