面向边缘设备的块扩散大型语言模型(LLM)的硬件加速
Hardware Acceleration of Block-Diffusion LLM for Edge Devices
浏览论文内容
中文总结 AI 辅助
针对边缘设备块扩散LLM推理的权重流量分摊与缓存问题,协同设计WIFiV-LPDDR、BRQ-KV、DAT-FFN组件,在Jetson级平台上实现显著能耗降低与延迟加速,且压缩模型性能损失极小。
中文摘要 AI 辅助
单流(单批次)边缘推理无法在多个请求间分摊权重流量。全注意力扩散LLM会在每一步重新计算整个序列;原生块扩散使已完成的块不可变且可精确缓存,但优化过程仍会流式传输前缀键值(KV)和前馈网络(FFN)权重。我们协同设计了WIFiV-LPDDR(一种用于精度标记读取的宽I/O LPDDR系统)、BRQ-KV(一种具有查询依赖逐条目精度的规范低秩加INT8残差前缀)和DAT-FFN(用于漂移映射规范替换、相邻阶段校正的低位增量或缓存状态保留,同时保持激活值未量化),三者均映射到输入 stationary 的混合精度 systolic 阵列。针对所评估的15亿/70亿参数模型在建模的Jetson级平台上,在报告的DAT-FFN设置下,该完整栈分别实现了3.79倍/3.96倍的算术平均能耗降低因子和2.88倍/4.44倍的算术平均延迟加速;每个对应的压缩模型基准分数较基线下降均小于1个绝对百分点。
英文摘要
Single-stream (batch-one) edge inference cannot amortize weight traffic across requests. Full-attention diffusion LLMs recompute the entire sequence at every step; native block diffusion makes completed blocks immutable and exactly cacheable, yet refinement still streams prefix KV and FFN weights. We co-design WIFiV-LPDDR, a wide-I/O LPDDR system for precision-tagged reads, BRQ-KV for a canonical low-rank-plus-INT8-residual prefix with query-dependent per-entry precision, and DAT-FFN for drift-mapped canonical replacement, adjacent-stage-corrected low-bit delta, or cached-state carry while keeping live activations unquantized. Both map to an input-stationary mixed-precision systolic array. For the evaluated 1.5B/7B models on modeled Jetson-class platforms, the full stack provides arithmetic-mean energy-reduction factors of 3.79x/3.96x and arithmetic-mean latency speedups of 2.88x/4.44x at the reported DAT-FFN settings; every corresponding compressed model-benchmark score drops by less than one absolute percentage point from its baseline.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。