AI 中文总结
该研究针对存储受限的大语言模型解码,提出窗口式存储屋顶线与双预算架构,发现降低每 token 字节数可提升边缘设备上的 Qwen3-30B-A3B 解码速度,专家预取无法缓解带宽饱和问题。
AI 中文摘要
在廉价硬件上的自回归解码并非受浮点运算(FLOPs)限制,而是受每个生成 token 必须在内存层次结构最慢的已填充层级间移动的字节数限制。我们将每 token 字节数视为一级设计轴,按地址确定性分类法组织,该分类法在 token 的前向传播期间,根据参数的获取地址何时可知对其进行分类:A0 类:在 token 采样时可知;A1 类:在注意力机制前可知;A2 类:按层依赖数据可知;A3 类:始终需要读取。这将预取调度简化为具有释放时间的单机可行性,从而得到闭式形式的窗口式屋顶线。我们在三个低于 1 亿参数的规模上运行双预算(每 token 字节数 × 存储)消融实验,随后将该框架应用于真实的大型混合专家(MoE)部署,并报告了屋顶线预测的重大负面结果:在运行 Qwen3-30B-A3B(4 比特,18GB)的 8GB 边缘板上,模型会溢出随机存取存储器(RAM),解码被固定在嵌入式多媒体卡(eMMC)带宽上限;预测性专家预取无济于事——无论是时间局部性预取(净负面效果),还是具有完美预测的轨迹驱动神谕,都无济于事,因为绑定约束是饱和总线上的字节量,而预取无法减少该字节量。有效的手段是降低每 token 字节数,直到模型适配快速层级:量化后适配 16GB 统一内存设备的同一模型,以 11.5 token/秒的速度在 GPU 上运行(速度提升 22 倍)。我们将此与 GPU 服务的专家预取预测器相协调:冻结模型探测以 91.2% 的准确率从注意力前状态预测 Qwen3-30B 的路由,这是一种与规模无关的可预测性属性,但仅当快速层级缓存了大部分模型且每 token 传输与计算相当的情况下,该属性才能转化为吞吐量——在 A100 外设组件互连高速(PCIe)卸载路径上已验证此情况成立,而在带宽受限的边缘存储上则不成立。可预测性并非速度提升;我们绘制了差距缩小的场景。
英文摘要
Autoregressive decoding on cheap hardware is bound not by FLOPs but by the bytes each generated token must move across the slowest populated tier of a memory hierarchy. We treat bytes-per-token as a first-class design axis, organized by an address-determinism taxonomy that classifies parameters by when their fetch address becomes known during a token's forward pass (A0: at token sampling; A1: before attention; A2: layerwise data-dependent; A3: always read). This reduces prefetch scheduling to single-machine feasibility with release times, yielding a closed-form windowed roofline. We run dual-budget (bytes-per-token times storage) ablations across three sub-100M scales, then take the framework to real large-MoE deployment and report a substantial negative result the roofline predicts: on an 8GB edge board running Qwen3-30B-A3B (4-bit, 18GB), the model overflows RAM and decode is pinned at the eMMC bandwidth ceiling; predictive expert prefetch does not help -- not temporal-locality prefetch (net-negative), not even a trace-driven oracle with perfect prediction -- because the binding constraint is byte volume over a saturated bus, which prefetch cannot reduce. The lever that works is reducing bytes-per-token until the model fits the fast tier: quantized to fit a 16GB unified-memory device, the same model runs GPU-resident at 11.5 tok/s (22x). We reconcile this with GPU-serving expert-prefetch predictors: a frozen-model probe predicts Qwen3-30B routing from the pre-attention state at 91.2%, a scale-invariant predictability property, but this converts to throughput only where the fast tier caches most of the model and per-token transfer is comparable to compute -- measured to hold on an A100 PCIe-offload path and to fail on bandwidth-walled edge storage. Predictability is not speedup; we chart where the gap closes.
CommentsWorkshop draft. Deployment numbers are single-run per configuration; negatives (prefetch/oracle) reported honestly. Code, data, and paper source: https://github.com/William2333ZZ/budgeting-bytes