发表机构
Florida State University(佛罗里达州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出轨迹局部自适应检索(TLAR)方法,利用模型自身生成轨迹作为草稿来源,结合检索与验证机制,提升长链推理中的解码效率与吞吐量。
AI 中文摘要
长链式思维推理增加了顺序解码的成本,同时产生了不断增长的潜在可重用延续历史。我们研究了这段历史何时能提供有用的草稿,并补充现有的起草者。受控来源比较揭示了轨迹特定的重用,这激发了我们的方法——轨迹局部自适应检索(TLAR)。TLAR从当前轨迹中检索近似匹配的延续,并利用最近的验证结果来调整检索激活和候选宽度。TLAR将检索到的延续与模型生成的草稿结合在一个共享的候选树中,通过精确验证保持目标模型的输出分布。在代码调试、数学和开放式写作中,我们的评估将来源重用、增量接受和执行成本联系起来。将TLAR与强检索基线相结合,在匹配的验证预算下提高了令牌接受率,并相对于草稿模型基线增加了端到端吞吐量。这些发现支持将生成的轨迹作为自适应推理的运行时记忆。
英文摘要
Long chain-of-thought reasoning increases sequential decoding cost while creating a growing history of potentially reusable continuations. We investigate when this history supplies useful drafts and complements an existing drafter. Controlled source comparisons reveal trajectory-specific reuse, motivating our method Trajectory-Local Adaptive Retrieval (TLAR). TLAR retrieves approximately matched continuations from the current trajectory and uses recent verification outcomes to adapt retrieval activation and candidate width. TLAR combines retrieved continuations with model-generated drafts in a shared candidate tree, preserving the target model's output distribution through exact verification. Across code debugging, mathematics, and open-ended writing, our evaluation connects source reuse, incremental acceptance, and execution cost. Combining TLAR with strong retrieval baselines improves token acceptance under matched verification budgets and increases end-to-end throughput over the draft-model baseline. These findings support generated trajectories as runtime memory for adaptive inference.
Comments33 pages