arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12114cs.OS

引入开销:在张量框架中采用文件支持的权重

The Ingestion Tax: Adopting File-Backed Weights in Tensor Frameworks

Yuan Si, Yufeng Lin, Daming Li, Jialu Zhang

中文总结 AI 辅助

该研究针对张量框架中权重引入开销问题,提出文件支持的权重采用方案,通过三部分执行契约消除开销,在多平台上实现性能提升,且支持多进程共享映射副本以降低内存开销。

中文摘要 AI 辅助

开放权重模型可占据中等容量区间:活跃权重作为缓存文件页适配DRAM,但框架所有的第二种表示形式要么无法适配,要么需在层运行时重新填充,因此小批量解码会为每个令牌重新读取权重。在集成内存和连贯内存系统上,检查点的文件页已占据GPU可读域,但加速器加载路径仍会将它们复制到框架所有的分配空间中。我们将此称为“引入开销”:操作系统以干净、可驱逐的文件页形式持有字节,架构使这些字节可被GPU读取,而只有框架的所有权模型阻碍了两者的直接关联。我们提出文件支持的权重采用方案:一个与框架无关的生产者通过MAP_SHARED映射每个张量,将页包装为无复制的GPU缓冲区,并导出一个DLPack胶囊,供PyTorch或MLX作为普通存储导入。仅零复制导入还不够:一项随机析因实验确立了三部分执行契约——读取映射、保持激活驻留在加速器上、在GPU上排序;若添加扩展放弃后两项,运行速度比基准慢2.3倍。在该契约下,采用此方案可消除开销且无速率损失:达到516 GB/s,而默认构造函数仅达53-82 GB/s,与驻留存储上的同一内核性能相当(差值范围[-0.66%, +0.48%];Qwen2.5-72B的令牌生成速度为7.14 vs 7.23 tok/s)。这种性能对等带来了所有权优势:N个进程解码一个映射副本,而驻留加载需支付N倍开销(在容量受限情况下,速度为5.5 vs 0.08 tok/s);65 GB检查点的首令牌生成速度提升6.4倍;Kimi K3主干阶段的每令牌时间从2.62 s降至0.35 s,较仅从存储读取提升3.8倍。该机制在AMD APU上以一半内存占用实现1.21倍的性能提升,在容量超限的GH200上与重叠流式处理性能相当,在PCIe上则损失39倍性能——字节路径由内存拓扑而非API决定。该部署规则将页缓存视为一等可回收的加速器存储层。

英文摘要

Open-weight models can occupy a middle capacity regime: active weights fit in DRAM as cached file pages, but a second framework-owned copy does not fit or must be refilled as layers run, so low-batch decode rereads the weights every token. On integrated and coherent-memory systems those file pages are already GPU-readable, yet ordinary loading paths copy them into framework allocations before use. We call this copy the ingestion tax. We present file-backed weight adoption: a framework-independent producer maps each tensor with MAP_SHARED, wraps the pages as a no-copy GPU buffer, and exports a DLPack capsule that PyTorch or MLX imports as ordinary storage. Zero-copy import alone is insufficient: the implementation must also keep activations accelerator-resident and establish ordering on the GPU; an adopter that omits both runs a dense decode stage 2.3x slower than stock in the live system. With both in place, adoption removes the tax: the public route reaches 516 GB/s versus 53-82 for the default constructors, matches the identical kernel over resident storage ([-0.66%, +0.48%], paired), and is within 1.3% of a resident control on a matched Qwen2.5-72B (7.14 vs. 7.23 tok/s). At the same throughput, the weights remain clean, shared, evictable file pages: N processes decode from one mapped copy where resident loading creates N copies (at capacity, 5.5 vs. 0.08 tok/s), and a 65 GB checkpoint cuts time to first token by 6.4x versus stock loading. In Kimi K3, a 2.8T-parameter MoE, the dense int8 spine stage falls from 2.62 to 0.35 s per token (7.5x; 3.8x from storage alone). The same mechanism improves llama.cpp by 1.21x at half the footprint on an AMD APU, falls inside the 5% selection band of overlapped streaming on a capacity-exceeding GH200 workload, and is 39x slower across PCIe. The deployment rule follows memory topology: adopt file pages only where the GPU can already read them.

↑