arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PI-Mem:通过并行迭代式记忆将长上下文推理推向360万 tokens

PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory

Dawei Liu, Haixu Song, Shuang Cheng, Shijie Wang, Haozheng Hou, Kaifeng Liu, Ermo Hua, Zhonghang Yuan, Zhijie Zhong, Yuchen Fan, Biqing Qi, Bowen Zhou

arXiv 2608.03048首次发表:更新:

AI 中文总结

该研究提出PI-Mem并行迭代记忆机制,在360万 tokens长上下文下,使Qwen系列模型在HotpotQA基准上较循环记忆基线实现准确率提升与推理加速,打破长上下文推理的准确率-效率权衡。

AI 中文摘要

长上下文推理仍是大型语言模型的关键瓶颈,因为近期的循环记忆方法面临两个固有挑战:分块顺序更新会用后续无关内容覆盖早期关键证据,且分块间的串行依赖限制了并行性,导致延迟随上下文长度增加而上升。为解决这些问题,我们提出PI-Mem(Parallel-Iterative Memory,并行迭代式记忆),这是一种在有限轮次内并行处理所有分块并迭代优化共享记忆的机制。每一轮中,PI-Mem基于当前记忆并行读取所有分块,从每个分块中选择新的或补充性证据,并将所选证据合并为紧凑的共享记忆供下一轮使用。为抑制冗余轮次,我们通过带有辅助轮次效率奖励的强化学习优化工作流程,使模型在积累足够证据后能自适应退出。我们在HotpotQA基准上使用Qwen3.5-35B-A3B和Qwen2.5-7B评估PI-Mem,上下文长度最高达360万 tokens,发现其相比循环记忆基线分别提升了+6.25和+7.81的绝对得分,同时实现了6.1倍和2.1倍的推理加速。这些结果表明,PI-Mem打破了长上下文推理中的准确率-效率权衡,为处理超长文档上的复杂多跳问答提供了可扩展的方法。

英文摘要

Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irrelevant content, and serial inter-chunk dependencies limit parallelism and cause latency to increase with context length. To address these issues, we propose PI-Mem (Parallel-Iterative Memory), a mechanism that processes all chunks in parallel and iteratively refines a shared memory over a bounded number of turns. In each turn, PI-Mem reads all chunks in parallel conditioned on the current memory, selects new or complementary evidence from each chunk, and merges the selected evidence into a compact shared memory for the next turn. To discourage redundant turns, we optimize the workflow through reinforcement learning with an auxiliary turn-efficiency reward, enabling the model to adaptively exit once sufficient evidence has been accumulated. We evaluate PI-Mem with Qwen3.5-35B-A3B and Qwen2.5-7B on the HotpotQA benchmark across context lengths up to 3.6 million tokens and find that it outperforms the recurrent-memory baseline by +6.25 and +7.81 absolute points while achieving 6.1$\times$ and 2.1$\times$ inference speedups, respectively. These results demonstrate that PI-Mem breaks the accuracy--efficiency trade-off in long-context reasoning and provides a scalable approach to complex multi-hop question answering over extremely long documents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑