发表机构
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM长上下文推理的挑战,ConvMem提出无需训练的层次卷积框架,将推理路径缩短为对数树,实现并行化,并在RULER基准上优于基线,避免过拟合。
AI 中文摘要
尽管大型语言模型(LLMs)已展现出令人印象深刻的能力,但由于固定的上下文限制,它们常常在处理极长上下文时遇到困难。为了解决这一问题,诸如MemAgent之类的顺序方法通过分段读取文本并迭代更新固定大小的记忆来扩展有效上下文。然而,这种顺序范式存在高延迟问题,并且需要昂贵的强化学习(RL)训练,这可能导致对特定数据集的过拟合。为了克服这些限制,我们提出了ConvMem,一个无需训练、高度可并行化的框架,将长上下文推理重新表述为层次卷积。受CNN启发,ConvMem将带有特定查询提示的LLM视为卷积核。该核层次化地总结文本片段,将推理路径从线性链缩短为对数树。具体来说,ConvMem集成了可配置步长和跳跃连接,以确保稳健的证据捕获和传播,同时采用多核卷积将复杂查询分解为解耦的语义通道。这种设计不仅减轻了错误累积,还实现了跨文本片段和推理线程的大规模并行化。在RULER-HotpotQA和RULER-2WikiMultiHopQA上的实验表明,ConvMem优于无需训练基线的性能,并避免了在分布外任务中经常观察到的RL训练模型对参数先验过拟合的风险。
英文摘要
While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and iteratively updating a fixed-size memory. However, this sequential paradigm suffers from high latency and requires costly reinforcement learning (RL) training, which can lead to overfitting on specific datasets. To overcome these limitations, we propose ConvMem, a training-free, highly parallelizable framework that reformulates long-context reasoning as a hierarchical convolution. Inspired by CNNs, ConvMem treats an LLM prompted with a specific query as a convolutional kernel. This kernel summarizes text segments hierarchically, shortening the reasoning path from a linear chain into a logarithmic tree. Specifically, ConvMem integrates \textit{Configurable Strides} and \textit{Skip Connections} to ensure robust evidence capture and propagation, while employing \textit{Multi-Kernel Convolution} to decompose complex queries into disentangled semantic channels. This design not only mitigates error accumulation but also enables massive parallelization across both text segments and reasoning threads. Experiments on RULER-HotpotQA and RULER-2WikiMultiHopQA demonstrate that ConvMem outperforms training-free baselines and avoids the risk of overfitting to parametric priors often observed in RL-trained models on out-of-distribution tasks.