ThinkRetrieve:用于测试时缩放的检索增强推理轨迹
ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling
浏览论文内容
中文总结 AI 辅助
针对大型推理模型测试时缩放的负面效果,提出ThinkRetrieve框架,通过每步动态检索已解决示例注入推理轨迹,在四个数据集上对五个模型实现最高60%的相对准确率提升。
中文摘要 AI 辅助
大型推理模型(LRMs)通过分配额外的推理时计算资源生成扩展的思维链推理来提升性能。但近期研究显示,顺序测试时缩放常产生递减甚至负面效果,因为更长的推理轨迹会出现不确定性增加、错误累积及偏离原始问题的情况。我们提出ThinkRetrieve,一种测试时缩放框架,它在LRMs的推理轨迹中,于每一步推理时动态检索已解决的示例。给定包含问题及分步解答的外部语料库,ThinkRetrieve在每个中间步骤检索相关范例并直接注入思维轨迹,为模型提供推理方式的指导,而非仅相关事实。在GSM-8K、MATH-500、AIME 2025和SciQ四个数据集上,针对五个参数规模为15亿至80亿的推理模型开展的实验表明,ThinkRetrieve相较标准测试时缩放能持续提升准确率,在AIME 2025上的相对提升最高达60%。
英文摘要
Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies reveal that sequential test-time scaling often yields diminishing or even negative returns, as longer traces exhibit increased uncertainty, error compounding, and drift from the original problem. We propose ThinkRetrieve, a test-time scaling framework that augments the reasoning traces of LRMs with dynamically retrieved solved examples at each reasoning step. Given an external corpus of problems paired with step-by-step solutions, ThinkRetrieve retrieves relevant exemplars at each intermediate step and injects them directly into the thinking trace, providing the model with guidance on how to reason rather than merely what facts are relevant. Experiments across five reasoning models (1.5B--8B parameters) on GSM-8K, MATH-500, AIME 2025, and SciQ demonstrate that ThinkRetrieve consistently improves accuracy over standard test-time scaling, with relative gains of up to $60\%$ on AIME 2025.
发表机构
- IIT Bombay(印度理工学院孟买分校)
- University of Maryland, College Park(马里兰大学帕克分校)
- Adobe Research(奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。