一个Token的成本是什么?一种基于多智能体混合的充分单Token计算量度量
What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute
- Duke University(杜克大学)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文通过多智能体混合方法测量单个token的充分计算量,发现大部分token计算冗余,并利用该映射优化模型路由和起草,显著降低延迟并保持准确率。
中文摘要 AI 辅助
大型语言模型对其生成的每个token都花费相同的计算量,无论每个token的生成难度如何。投机解码和模型路由等方法基于这样一个前提:大部分计算是不必要的,然而单个token实际所需的计算量尚未被测量。我们通过多智能体混合(MoA)的视角来测量它:一个由来自三个家族的十五个容量递增的语言模型组成的小组,其中每个智能体在给定正确的前序token条件下,逐token尝试复现一个参考序列。我们将成功的最小智能体的推理成本定义为该token的充分计算量,这为token所需计算量提供了一个上界。在三个核心基准上,一个0.5B的智能体能够复现92%至95%的参考token。在Qwen、OLMo和R1蒸馏小组中,最昂贵的10%的token占估计FLOPs的64%至80%。在所有500个MATH-500问题上,由MoA导出的映射帮助模型路由将预计延迟从7.59秒降低到5.12秒,同时相对于最佳置信度路由基线,准确率略有提升。MoA映射帮助起草过程比固定窗口起草在相似准确率下使用少32.6%的起草token,并降低约20%的预计延迟。这些比较揭示了剩余的资源分配空间,激励了利用充分计算量结构的控制器。
英文摘要
Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually requires has not been measured. We measure it through a Mixture-of-Agents (MoA) lens: a panel of fifteen language models of increasing capacity, drawn from three families, in which every agent attempts to reproduce a reference sequence token by token, conditioned on the correct preceding tokens. We define the inference cost of the smallest agent that succeeds as the token's sufficient compute, which upper-bounds what the token requires. On three core benchmarks, a 0.5B agent reproduces 92--95\% of reference tokens. Across Qwen, OLMo, and R1-distilled panels, the most expensive 10\% account for 64--80\% of estimated FLOPs. On all 500 MATH-500 problems, the MoA-derived map helps model routing reduce projected latency from 7.59 to 5.12 seconds while slightly improving accuracy, relative to the best confidence-routing baseline. The MoA-map helps drafting use 32.6\% fewer draft tokens and approximately 20\% lower projected latency than fixed-window drafting at similar accuracy. These comparisons reveal remaining allocation headroom, motivating controllers that exploit sufficient-compute structure.