Jet-Long:使用动态双焦点旋转位置编码的高效长上下文扩展
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对现代语言模型长上下文应用中零样本上下文扩展问题,提出Jet-Long方法,通过动态双焦点RoPE及相关技术,在推理时开销小,在多种模型和基准测试中表现优异,还能推广到其他架构且超参数弹性强。
AI中文摘要:
现代语言模型越来越多地应用于长上下文应用,如检索增强生成、存储库级编码和智能体工作流程,这使得零样本上下文扩展成为开放权重检查点的主要部署路径。现有零样本方法预先固定单个缩放因子,存在问题。本文提出Jet-Long,一种免调优的零样本方法,它将局部忠实于旋转位置编码(RoPE)的窗口与远程窗口配对,缩放因子动态适应序列长度。还介绍了包含-排除注意力合并和动态RoPE校正旋转等。实验表明,在H100上长上下文预填充吞吐量可达1.39倍FA2,单批次生成开销≤4%。在Qwen模型及相关基准测试中表现出色,还可推广到混合注意力架构,且超参数弹性强,便于部署。
英文摘要:
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. The dominant zero-shot methods (YaRN, Self-Extend, DCA) fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity while a conservative one breaks down at long contexts; recent length-aware variants adapt the mapping, but with a fitted or distance-dependent schedule. We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically to the current sequence length via a parameter-free analytic schedule, recovering the base model exactly at short inputs while extrapolating cleanly at long ones. An inclusion-exclusion attention merge and on-the-fly RoPE correction enable a fused CuTe implementation. On H100 at 64K-128K, prefill retains 83-88% of FlashAttention-3 throughput across the evaluated Qwen3 sizes and 88-93% of a matched CuTe control; Qwen3-8B single-batch generation reaches 1.04-1.08 times FlashAttention-3 throughput. On Qwen3-1.7B/4B/8B up to 128K context, Jet-Long leads RULER by +4.79/+2.18/+2.03 percentage points over the strongest baseline at 1.7B/4B/8B, achieves the best overall accuracy on HELMET-RAG (a benchmark identified by HELMET as the most efficient predictor of downstream long-context performance) and attains the lowest PG-19 perplexity. Additional evaluations cover Meta-Llama-3-8B, post-trained Qwen3 checkpoints, and the hybrid Jet-Nemotron architecture, supporting broader applicability without retraining. The local-window hyperparameter remains robust across the tested settings.