发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文以Kimi K3模型为对象,探究智能体编码任务中任务规格对token消耗的影响,发现简化任务规格会提升token消耗,拟合的预测器可较优地预测token消耗,为评估AI编码工作流成本提供方法。
AI 中文摘要
智能体编码工作流现已广泛部署于实际系统中。借助长程推理与工具使用,token用量已成为成本与效率层面的重要考量。两位使用AI的工程师会以不同方式解决同一问题。任务规格如何影响智能体的token消耗,以及该消耗是否可提前预测,仍是未解决的问题。本文研究了不同任务规格对智能体token消耗的影响,采用Kimi K3模型并设置三种思考力度。在2700次运行中,我们发现将完整任务规格简化为最简用户故事会使token消耗提升29.7%,而运行间方差不受任何提示更改的影响。我们还发现提示敏感性具有任务依赖性,范围为13%至115%。我们拟合了一个简单预测器,该预测器可通过对未见过的任务进行一次低成本探测,在36%的误差内对完整任务规格和思考力度配置的token消耗分布进行定价,优于先前预测token消耗的工作。本文研究提供了量化任务规格对智能体token消耗影响的初步结果,并引入了一种可用于系统评估AI编码工作流成本的方法。
英文摘要
Agentic coding workflows are now widely deployed in real-world systems. With long-horizon reasoning and tool use, token usage has become an important consideration for both cost and efficiency. Two engineers using AI will solve the same problem differently. How the specification of a task shapes an agent's token spend, and whether that spend can be predicted in advance, are open questions. Here, we study the effects of different task specifications on agentic token spend with the Kimi K3 model at three thinking efforts. Across $2,700$ runs, we show that reducing a full task specification to a bare user story raises token spend by $29.7\%$, while run-to-run variance remains unaffected by any prompt changes. We show that prompt-sensitivity is task-dependent, running from $13\%$ to $115\%$. We fit a simple predictor that can price a full distribution of task specifications and thinking effort configurations from a single cheap probe on an unseen task within $36\%$, improving over prior work in predicting token spend. Our work provides initial results quantifying the effects of task specification on agentic token spend and introduces a method that can be used to systematically evaluate the cost of AI coding workflows.