arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

记忆的代价:一种经过校准的计算能量定律

The Price of Remembering: A Calibrated Energy Law for Computation

Mohamed Amine Bergach

arXiv 2609.00744首次发表:更新:

发表机构

Illumina(因美纳)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种包含计算操作、数据存储租金与移动路费的校准计算能量定律,推导得出精确注意力能量随上下文长度平方增长等结论,揭示长上下文服务的带宽限制问题及内存层级的效率优势。

AI 中文摘要

计算机消耗的大部分能量并非用于计算,而是用于维持数据:存储在快速存储器中的每一位在保留期间每秒都会消耗功率,且每在存储层级间移动一次就会再次消耗能量。我们将第一种成本称为“租金”,第二种成本称为“路费”,并提出一条定律:一次计算的能量至少等于其操作本身的能量,加上每一位存活期间的租金,再加上每一位移动的路费。该定律背后的模型也对控制过程定价:不存在免费的时钟,任何未定价的寄存器都会使定理失效。一个引理支撑了这些结果:对一个值的每一次使用都需通过租金、路费或重新计算来支付。由此产生三个结论:精确注意力机制会为每个新 token 召回所有过往 token,其能量随上下文长度的平方增长,而固定状态循环模型的能量呈线性增长;该平方关系是永不重读过往 token 的机器的定理,且在给定的服务假设下,这一平方值对应每个过往 token 的路费,在接近 10^4 个 token 时会超过循环模型自身的算术成本,此时长上下文服务如今受带宽限制。累积内存边界成为焦耳下限:在大多数输入下,对拥有易失性工作存储的任何顺序机器,排序 n 个元素的租金与 n²/log n 位步长成比例,scrypt 的边界使得每次密码猜测的焦耳成本无法通过并行计算降低。该定律经过校准:在一个存储可物理移动的合成 45 nm 处理器上,门级测得的租金常数为每周期每个占用插槽 0.82 pJ,加上每个净跃迁 3.4 fJ,在操作隔离硬件的每一次旋转运行中误差在 4% 以内。该租金的五分之四来自时钟:维持数据主要是确定何时进行操作。内存层级的存在是因为分层存储比扁平化存储更高效,在我们的仪器上效率为 83 倍。我们还说明了该定律如何被证伪。

英文摘要

Where does a computer's energy go? Mostly into keeping, not into computing. A bit held in fast storage draws power for every second it stays there, and it costs energy again each time it moves between storage levels. We call the first cost \emph{rent} and the second \emph{fare}, and we state one law: the energy of a computation is at least its operations, plus rent on every live bit for as long as it lives, plus fare on every bit moved. The model under the law prices control as well as data. There is no free clock, and any unpriced register would make the theorems false. One lemma does most of the work: every use of a value is paid for by rent, by fare, or by computing the value again. Three things follow. Exact attention brings every past token back for every new one, so its energy grows with the square of the context length, while a recurrent model with a fixed state grows linearly. The square is a theorem for machines that never re-read past tokens. Under a stated serving hypothesis it is the fare on every past token, which passes the model's own arithmetic near ten thousand tokens, the point where long-context serving becomes bandwidth-bound today. Known bounds on memory over time become joule floors: on any sequential machine with volatile working storage, sorting $n$ items pays rent proportional to $n^2/\log n$ bit-steps on most inputs, and the bound for scrypt makes every password guess cost joules that no amount of parallel hardware reduces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑