arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28919cs.AIcs.CR

控制缰绳,控制成本:企业环境中 AI 编码代理的路由与治理

Harness Tokenomics: A Router for the Enterprise Agentic Control Plane

Ted Kwartler, Alan Aqrawi, Arian Abbasi

首次发表
浏览论文内容

中文总结 AI 辅助

针对企业AI编码代理成本失控问题,提出基于校准分类器Jev的可定制路由方法,在会话边界智能切换模型,可回收14%-21%模型支出,并给出治理风险的控制平面方案。

中文摘要 AI 辅助

运行 AI 编码代理的产品(即“缰绳”)正在激增,企业正将其推广给员工:最初只有几百个席位的试点项目,如今已扩展至数万个席位。大多数企业并不自行构建这些缰绳,而是从大型供应商处购买,例如 Anthropic 的 Claude Code 或 OpenAI 的 Codex。缰绳决定哪个模型回答问题、模型读取什么内容、如何使用提示缓存以及运行哪些子代理,因此它决定了价格表上的费率,并设定了在该费率下的购买量。企业若保留专有或未经调优的缰绳并采用其默认设置,则继承了这些选择及其账单。我们构建了一个快速、可定制的路由器,其中 Jev(一种具有校准概率的分类器)根据企业自带的代理请求分类法为每个提示打上标签。由于一个用户轮次涉及多个请求,而这些请求共享一个属于单个模型的提示缓存,因此路由器仅在无需重建现有对话缓存的情况下移动工作负载:即在会话开始时、侧通道中以及子代理启动时。根据价格表,我们推导出任务中途切换何时能收回成本,以及一个交叉点:在长时间、工具密集的会话中,价格最高的模型比次一级模型成本更低,这一点通过从公共数据集重新定价约 10,000 个真实会话得到了证实。在一个模拟的 10,000 席位企业中,用户行为取自这些数据集,路由器在 Anthropic 2026 年 9 月 21 日的标价下,可回收模型支出的 14% 至 21%,即每年 330 万至 500 万美元。本文还绘制了二十个缰绳的风险图谱,对依赖单一供应商模型的价格进行了评估,并提出了一种企业可从内部立即运行的控制平面,同时提供了一个阶梯式决策框架,用于后续决定是否自行拥有缰绳。

英文摘要

Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees. What started as pilots with a few hundred seats is now scaling to tens of thousands. Enterprises rarely build these harnesses and usually adopt the ones model vendors bundle with their seats and APIs, such as Anthropic's Claude Code or OpenAI's Codex. A harness decides which model answers, what the model reads, how the prompt cache is used and which subagents run, so it picks the rate on the price sheet and sets the volume bought at it. Enterprises that keep a proprietary or untuned harness at its defaults inherit these choices and their bill. We build a fast, customisable router in which Jev, a classifier with calibrated probabilities, labels every prompt against a bring-your-own taxonomy of agentic requests. One user turn is many requests over a prompt cache that belongs to one model, so the router moves work only where no running conversation has to rebuild its cache. It routes at session start, in side lanes and at subagent launch. From the price sheet we derive when a mid-task switch pays back. We also find a crossover, in which the highest-priced model costs less than the next tier on long tool-heavy sessions, and repricing about 10,000 real sessions from public datasets confirms it. In an emulated enterprise of 10,000 seats with user behaviour taken from these datasets, the router recovers 13 to 21% of model spend at Anthropic's list prices of 21 September 2026, \$3.3M to \$5.1M a year. The paper also maps the risks across twenty-one harnesses, prices the dependence on one vendor's models, and proposes a control plane that enterprises can run from within, starting now, with a ladder for deciding later whether to own the harness.

发表机构

  • Accenture Responsible AI(埃森哲负责任人工智能)
  • Accenture Americas Advanced AI Practice(埃森哲美洲高级人工智能实践)
  • Harvard Extension School, Harvard University(哈佛大学继续教育学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑