发表机构
Alibaba Group; Tsinghua University(阿里巴巴集团; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对有限上下文长程推理的冗余、溢出与过早猜测问题,提出可学习中间接口方法ThinkReset,通过接口写回与重置优化,在多基准测试中提升了固定窗口下的推理成功率。
AI 中文摘要
长思维链推理可提升复杂问题解决性能,但会引入冗余累积、上下文溢出与错误锚定问题。我们认为,在有限上下文窗口下,核心瓶颈并非轨迹压缩或测试时控制,而是缺乏可替代被丢弃历史、支持持续求解的可复用中间接口。我们进一步指出结果奖励驱动的长链强化学习的关键失效模式:当模型在窗口接近耗尽前未解决任务时,最终答案奖励会促使模型过早猜测而非继续仔细推理。我们提出ThinkReset,即该观点的文本空间实例化,其通过接口写回与重置显式构建可复用中间接口,并直接优化重置后的推理成功率。在多个长程推理基准测试中,该方法在固定上下文窗口下持续提升了任务成功率。
英文摘要
Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is not trajectory compression or test-time control, but the absence of a reusable intermediate interface that can replace discarded history and support continued solving. We further identify a key failure mode of outcome-reward-driven long-chain reinforcement learning: when the model has not solved the task before the window is nearly exhausted, the final-answer reward encourages premature guessing rather than continued careful reasoning. We propose ThinkReset, a text-space instantiation of this view. ThinkReset explicitly constructs reusable intermediate interfaces through interface writeback and reset, and directly optimizes post-reset continuation success. Across multiple long-horizon reasoning benchmarks, this perspective consistently improves success rates under fixed context windows.