arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21265cs.CL

记忆增强解锁高效思维链推理

Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen, Junyuan Shang, Tingwen Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出无需训练的记忆增强压缩框架,构建可复用推理记忆提升思维草稿压缩效果,在多任务上获显著准确率提升并加速推理,兼容多种压缩机制。

中文摘要 AI 辅助

大型语言模型常依赖思维链(CoT)推理解决复杂任务,但冗长的推理轨迹会带来大量推理开销。思维链压缩可缩短生成内容,但过度压缩可能破坏逻辑连贯性并降低性能。我们将这种权衡形式化为上下文-生成替代定律,其中显式推理上下文替代部分解码时生成内容。基于该原理,我们提出记忆增强压缩(Memory-Augmented Compression),这是一种无需训练的框架,可从历史轨迹构建可复用的推理记忆,并将其作为预填充侧支架进行检索。这些记忆并非使用原始演示,而是总结可复用的推理模式、关键约束和核心操作,以补偿压缩过程中丢失的信息。实验表明,在数学推理、复杂推理和科学问答任务中,记忆始终提升基于提示的思维草稿(CoD)压缩效果,在GSM8K、MATH、BBH和MMLU-Sci上,相较于CoD分别获得21.4、28.0、29.5和6.61个百分点的准确率提升,同时比标准CoT实现1.14至1.49倍的推理延迟加速。记忆还与令牌级、推理轨迹级和推理状态级压缩机制兼容。进一步分析显示,性能提升源于相关的推理记忆,而非单纯增加上下文长度。

英文摘要

Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex tasks, but verbose reasoning traces introduce substantial inference overhead. CoT compression shortens generation, yet aggressive compression may disrupt logical coherence and degrade performance. We formalize this trade-off as the Context-Generation Substitution Law, where explicit reasoning context substitutes for part of decode-time generation. Based on this principle, we propose Memory-Augmented Compression, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds. Rather than using raw demonstrations, these memories summarize reusable reasoning patterns, key constraints, and critical operations to compensate for information lost during compression. Experiments show that Memory consistently improves prompt-based Chain-of-Draft (CoD) compression across mathematical reasoning, complex reasoning, and science question answering tasks, yielding accuracy gains of 21.4, 28.0, 29.5, and 6.61 points over CoD on GSM8K, MATH, BBH, and MMLU-Sci, while achieving a 1.14-1.49x latency speedup latency speedup over standard CoT. Memory is also compatible with token-level, reasoning-trace-level, and inference-state compression mechanisms.

发表机构

  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
  • School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
  • Baidu Inc.(百度公司)
  • Tencent Inc.(腾讯公司)

机构由 AI 辅助整理,请以论文原文为准。

↑