arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TileMix:面向大语言模型推理加速的以块为中心的混合精度注意力机制

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng

arXiv 2608.17336首次发表:更新:

发表机构

University of North Texas; Saint Louis University(北得克萨斯大学; 圣路易斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TileMix是一种以块为中心的混合精度注意力内核,通过硬件对齐的分数块路由实现LLM推理加速,可恢复长上下文质量并提升预填充吞吐量,支持多种模型与任务场景。

AI 中文摘要

大语言模型(LLM)的长上下文预填充会产生大量计算和内存流量,因为密集自注意力会计算二次方的查询-键分数。现有方法要么使用统一的低精度路径,要么选择标记交互,而在融合密集注意力中,硬件对齐的分数块上的空间精度路由仍未被利用。我们提出TileMix,一种以块为中心的精度路由内核,它将数值精度作为融合密集注意力内分数块组上的可执行空间决策。TileMix将注意力矩阵划分为硬件对齐的分数块,将路由决策打包为紧凑的位掩码,并通过FP16或INT8分数计算调度每个块组,同时两条路径更新共享的在线softmax状态。可扩展的精度分组允许每个路由位控制多个相邻键块,在长上下文下保留硬件对齐的计算块和紧凑元数据。通过对所有合法块组进行路由,TileMix保留了密集标记连接,无需训练,支持分组查询注意力、可变长度批次以及INT8键/值缓存。在LLaMA、Qwen和Vicuna上的LongEval、LV-Eval和A100预填充基准测试中,TileMix恢复了统一INT8下丢失的长上下文质量,并提高了FP16下的预填充吞吐量,在各模型系列中实现了可控的精度-效率权衡前沿。实现代码可在指定URL获取。

英文摘要

Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial precision routing over hardware-aligned score tiles outside fused dense attention. We introduce TileMix, a tile-centric precision-routing kernel that makes numerical precision an executable spatial decision over score-tile groups within fused dense attention. TileMix partitions the attention matrix into hardware-aligned score tiles, packs routing decisions into compact bitmasks, and dispatches each tile group through FP16 or INT8 score computation while both paths update a shared online-softmax state. Scalable precision grouping lets each routing bit govern multiple adjacent key tiles, preserving hardware-aligned compute tiles and compact metadata at long contexts. By routing all legal tile groups, TileMix preserves dense token connectivity, requires no training, and supports grouped-query attention, variable-length batches, and INT8 key/value caches. Across LongEval, LV-Eval, and A100 prefill benchmarks on LLaMA, Qwen, and Vicuna, TileMix recovers long-context quality lost under uniform INT8 and improves prefill throughput over FP16, yielding a controllable accuracy-efficiency frontier across model families. The implementation is available at https://github.com/HanzhiZhang-Ulrica/TileMix.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑