arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15105cs.AI

使用有限VRAM进行长上下文微调

Long-Context Fine-Tuning with Limited VRAM

  • BMW Group(宝马集团)

机构由 AI 辅助整理,请以论文原文为准。

Vladimir Fedosov, Aleksandr Sazhin, Artemiy Grinenko, Frank Woernle

AI总结:

研究针对长训练序列成本高的问题,结合分层全局注意力(HGA)、逐段反向传播和分层KV存储,实现有限VRAM下的长上下文微调,在训练长度和吞吐量上有提升,且HGA也可用于检索和生成,相关服务实现正开发。

AI中文摘要:

参数高效微调可减少模型和优化器内存,但密集注意力仍使长训练序列成本高昂。我们将分层全局注意力(HGA)与逐段反向传播和分层KV存储相结合。在VRAM中只有活动段保持可微;旧的KV被分离到RAM或NVMe中,HGA为每个查询块加载一组有界的精确历史令牌。在具有4位QLoRA和PG19的Qwen3-8B上,使用16GB Quadro RTX 5000进行密集训练时,2048个令牌可以适配,但4096个令牌时失败,而HGA在峰值VRAM为15.28GB时可达16384个令牌。在评估中,同一适配器在此卡上可处理131072个令牌;VRAM不是恒定的,而是随着常驻块摘要逐渐增长,因此RAM和NVMe容量设定了超出这些长度的实际限制。在共享的2K训练长度下,HGA训练和密集训练的适配器在相同的密集注意力读出下分别获得2.7405和2.7383奈特,而原始模型获得2.9541。在此边界处,HGA训练已经略快(217.75对207.02令牌/秒),并且HGA与密集训练的吞吐量比从1K提高到2K;由于HGA使每个令牌的关注历史集大致恒定,而每个令牌的密集工作量增加,我们预计随着上下文增长,这种领先优势会扩大。密集注意力用于主要质量和检索比较,以便它们测量学习到的权重并与标准生成框架兼容。HGA也可用于检索和生成;优化的生产级服务实现正在开发中。

英文摘要:

Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive. We combine Hierarchical Global Attention (HGA) with segment-wise backpropagation and tiered KV storage. Only the active segment remains differentiable in VRAM; older KV is detached into RAM or NVMe, and HGA loads a bounded set of exact historical tokens for each query block. On Qwen3-8B with 4-bit QLoRA and PG19, dense training on a 16 GB Quadro RTX 5000 fits 2,048 tokens but fails at 4,096, whereas HGA reaches 16,384 tokens with 15.28 GB peak VRAM. Under evaluation the same adapter runs through 131,072 tokens on this card; VRAM is not constant but grows gently with the resident chunk summaries, so RAM and NVMe capacity set the practical limit beyond these lengths. At the shared 2K training length, HGA-trained and dense-trained adapters obtain 2.7405 and 2.7383 nat under the same dense-attention readout, while the stock model obtains 2.9541. At this boundary HGA training is already marginally faster (217.75 vs. 207.02 tokens/s), and the HGA-to-dense throughput ratio improves from 1K to 2K; because HGA keeps the attended historical set per token approximately constant while dense work per token grows, we expect this lead to widen as context grows. Dense attention is used for the main quality and retrieval comparisons so that they measure the learned weights and remain compatible with standard generation frameworks. HGA can also be used for retrieval and generation; an optimized production-grade serving implementation is under development.

↑