arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProxyFormer:用于超长上下文和高分辨率生成的双流代理架构

ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation

Zhongpan Tang

arXiv 2608.23463首次发表:更新:

AI 中文总结

ProxyFormer通过双流代理架构缓解超长上下文与高分辨率生成的注意力二次增长瓶颈,可扩展可训练序列长度至0.7M,在多针检索任务中保持高准确率,初步验证了其在图像生成中的可行性。

AI 中文摘要

注意力计算与键值(KV)缓存随序列长度呈二次增长,这是超长上下文语言模型和高分辨率生成模型的核心瓶颈。我们提出ProxyFormer,一种基于代理token的通用双流架构。在每一层中,细粒度局部特征自下而上压缩为一小组代理状态;仅在压缩后的代理空间中执行开销高昂的全局交互;经全局上下文化的代理随后自上而下解压并注入回局部流。由于局部流在各层间持续存在,一次压缩步骤未捕获的细粒度信息可被后续优化使用,缓解了传统一次性压缩的不可逆信息损失。我们进一步引入因子化多级压缩/解压、层级动态压缩比、不对称双嵌入及仅代理KV缓存推理方案。在批次大小为1的16GB GPU上,标准仅解码器模型仅能训练约20K token的序列,而压缩比为64的ProxyFormer可将可训练序列长度扩展至约0.7M。在1,048,576 token的多针检索任务中,以64K窗口训练的模型保持92%-95%的检索准确率;以8K窗口训练的模型外推至256K token时准确率超过94%。初步图像生成实验证明了ProxyFormer对像素空间和潜在空间流匹配的可行性。

英文摘要

The quadratic growth of attention computation and key-value (KV) cache with respect to sequence length is a central bottleneck for ultra-long-context language models and high-resolution generative models. We propose ProxyFormer, a general dual-stream architecture built upon proxy tokens. In each layer, fine-grained local features are compressed bottom-up into a small set of proxy states; expensive global interactions are performed only in the compressed proxy space; the globally contextualized proxies are then decompressed and injected top-down back into the local stream. Because the local stream persists across layers, fine-grained information that is not captured by one compression step remains accessible for later refinement, alleviating the irreversible information loss of conventional one-shot compression. We further introduce factorized multi-level compression/decompression, layer-wise dynamic compression ratios, asymmetric dual embeddings, and a proxy-only KV-cache inference scheme. On a 16GB GPU with batch size 1, a standard decoder-only model can train sequences of only about 20K tokens, whereas ProxyFormer with a compression ratio of 64 extends the trainable sequence length to about 0.7M. A model trained with a 64K window retains 92%-95% retrieval accuracy on a multi-needle retrieval task with 1,048,576 tokens, and a model trained with an 8K window exceeds 94% accuracy when extrapolated to 256K tokens. Preliminary image-generation experiments demonstrate the feasibility of ProxyFormer for both pixel-space and latent-space flow matching.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑