arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

来自先前问题的计算能否帮助LLMs解决新问题?

Can Computation from Earlier Problems Help LLMs Solve New Ones?

Jipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song

arXiv 2609.39394首次发表:更新:

发表机构

Gaoling School of Artificial Intelligence, Renmin University of China; University of Modena and Reggio Emilia(中国人民大学高瓴人工智能学院; 摩德纳和雷焦艾米利亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨历史计算对LLM解决新问题的影响,提出STAIR方法,通过固定存储库复用先前键值并重定向查询,仅训练1.2万参数,在三个Qwen模型和四个基准上平均后续轮次准确率最高提升11.67个百分点。

AI 中文摘要

大型语言模型常常在同一对话中解决独立的问题。来自先前问题的计算能否帮助它们解决新问题?为回答此问题,我们首先进行初步实验,表明保留的历史记录可以提高或降低后续轮次的准确率,即使在同一领域内也是如此。为理解这些效应,我们使用受控重放来隔离每个问题-历史配对特有的内部状态变化。在不同历史中,这些变化保留了当前问题之间的相似关系。为在保留历史的情况下改进推理,我们引入STAIR(用于查询间复用的陈旧令牌注意力)。STAIR在固定存储库中捕获先前响应生成中的键和值。它学会在提示处理期间当当前查询读取该存储库时重定向它们。基础模型保持冻结;仅训练12,288个参数。在三个Qwen模型和四个基准上,STAIR相比带历史记录的未修改模型,将平均后续轮次准确率最多提升11.67个百分点。

英文摘要

Large language models often solve independent problems in the same conversation. Can computation from earlier problems help them solve new ones? To answer this question, we first conduct preliminary experiments showing that retained history can raise or lower later-turn accuracy, even within the same domain. To understand these effects, we use controlled replay to isolate internal state changes specific to each problem-history pairing. Across different histories, these changes preserve similar relationships among current problems. To improve reasoning under retained history, we introduce STAIR (Stale-Token Attention for Inter-query Reuse). STAIR captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.

Comments29 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑