arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

驾驭效应:编排设计如何设定企业智能体人工智能的代币经济学

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh

arXiv 2607.06906首次发表:更新:

发表机构

Writer, Inc.(作家公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究智能体人工智能开发中代币最大化问题及应对方法,核心方法是通过控制交换实验凸显驾驭层作用,主要贡献为证明驾驭层可降低成本、提升效率和质量,且效率提升适用于所有模型,质量提升与模型强度相关。

AI 中文摘要

当今智能体人工智能开发依赖代币最大化,即通过代币购买能力,导致每个任务的代币增长快于任务价值,尽管代币价格下降,但总支出仍上升。我们认为应对代币最大化的关键杠杆是驾驭层,它能组装上下文、暴露工具、安排轮次、委托工作并承载企业可观测性和治理。通过控制交换实验,固定模型仅改变编排层,结果显示驾驭层可降低任务混合成本41%、中位数运行时间44%、每个任务的代币数38%,且任务完成质量相当。效率提升对所有模型都适用,质量提升与模型基线强度相关,我们称之为驾驭杠杆。每美元质量提升82%,每百万代币的任务完成数增加。我们还形式化了编排层的代币经济学,详细阐述了六种机制,并比较了六个智能体系统,认为驾驭层是能提升组织运行所有模型效率的关键组件。

英文摘要

Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway. We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries enterprise observability and governance. We isolate it with a controlled swap: 22 locked evaluation tasks, six foundation models (Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, Palmyra X6), changing only the orchestration layer -- a frozen conventional production loop versus the Writer Agent Harness. Holding models constant, the harness cuts blended cost per task 41% ($0.21->$0.12), median wall-clock 44% (48s->27s), and tokens per task 38% (14.2k->8.8k), with task-completion quality at parity (0.78->0.81, directional at this sample size). Efficiency is model-invariant -- every model gets cheaper (33-61%) -- while quality gains are capability-dependent: a model's gain correlates almost perfectly with its baseline strength (r=0.99, n=6), a phenomenon we term harness leverage. Quality per dollar rises 82%; task-completions per million tokens rise from 54.9 to 92.0. On this workload the orchestration layer moved cost per task more than the full spread of the model menu did. We formalize token economics at the orchestration layer (including effective input price under prompt caching), detail the six mechanism families behind the effect -- cache-shape discipline to failure-spend governance -- compare six widely used agent systems on the same axes, and argue the harness is the one component whose efficiency multiplies across every model an organization runs -- present and future.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑