代理总成本:多智能体LLM工作流中记忆注入成本的精确归因
Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows
浏览论文内容
中文总结 AI 辅助
针对多智能体LLM工作流中记忆注入成本不可见的问题,提出TCA分解与精确归因方法,实证显示其占比可观且可通过调整检索窗口容量控制。
中文摘要 AI 辅助
在多智能体大语言模型(LLM)工作流中,每个节点都会从记忆中检索上下文并将其注入到其提示中,这些注入的令牌以与系统提示和用户查询相同的每令牌价格计为输入令牌。生产可观测性工具报告总令牌成本,但不区分节点生成的令牌与其接收的令牌,因此这部分账单对支付它的团队来说是不可见的。我们引入了代理总成本(TCA),将多智能体工作流成本分解为基础提示、推理、记忆注入、未命中惩罚和上下文累积组件,并提出一种精确归因方法:一种两遍、非计费令牌计数,直接测量注入的令牌,而不是通过词数代理进行估计。在针对真实模型API执行的200任务企业基准测试中,记忆注入占编译时优化器可操作的变量成本的13.6%,约占完整计费成本的12%,其份额从工作流深度为1时的结构性零上升到深度为6时的27.6%。在测量范围内,注入令牌随深度线性增长(R^2 = 0.9974,深度2至6);二次拟合产生负的前导系数,因此数据在这些深度不表现出凸增长。我们表明该组件在固定模型层级下可控:将检索窗口容量从32个条目减少到2个条目,注入令牌降低28.7%,精度变化在种子级变异范围内。我们完整报告,我们的图重写变换在隔离时近似成本中性,五个分解项中有两个在此测试台中构造上为零,总工作流成本由模型层级分配主导,我们将其固定并视为先前工作。未评估提示缓存;所有数字均针对未缓存情况。
英文摘要
Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same per-token price as the system prompt and the user query. Production observability tools report total token cost but do not separate the tokens a node generates from the tokens it is handed, so this component of the bill is invisible to the teams paying it. We introduce the Total Cost of Agency (TCA), a decomposition of multi-agent workflow cost into base prompt, inference, memory injection, miss penalty and context-accumulation components, and an exact attribution method: a two-pass, non-billable token count that measures injected tokens directly rather than estimating them from word-count proxies. On a 200-task enterprise benchmark executed against real model APIs, memory injection accounts for 13.6 percent of the variable cost a compile-time optimizer can act on, about 12 percent of the full billed cost, and its share rises from a structural zero at workflow depth one to 27.6 percent at depth six. Injected tokens grow linearly with depth over the measured range (R^2 = 0.9974, depths two through six); a quadratic fit yields a negative leading coefficient, so the data do not exhibit convex growth at these depths. We show the component is controllable at fixed model tier: reducing the retrieval window capacity from 32 to 2 entries lowers injected tokens by 28.7 percent with an accuracy change within seed-level variation. We report in full that our graph-rewriting transforms are approximately cost-neutral in isolation, that two of the five decomposition terms are zero by construction in this harness, and that total workflow cost is dominated by model tier assignment, which we hold fixed and treat as prior work. Prompt caching is not evaluated; all figures are for the uncached case.