发表机构
Université de Montréal; Mila – Quebec AI Institute(蒙特利尔大学; 米拉-魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LatentHarness通过反事实策略蒸馏统一记忆与潜在推理,在长上下文基准上相对提升2.8%和10.0%,速度提升5.9倍。
AI 中文摘要
长上下文推理面临两个互补的瓶颈:在长输入中保留证据,以及在众多推理步骤中维持计算。现有方法大多分别处理这两个问题,外部记忆扩展了对遥远证据的访问,而潜在推理压缩了多步计算。我们引入了LatentHarness,它将记忆访问和潜在推理统一为顺序的潜在动作选择。在每个内部步骤中,模型选择THINK进行进一步计算,从输入证据和中间推理状态的快速权重记忆中进行RECALL,或EXIT以发出下一个词元。我们使用反事实策略蒸馏来训练该策略,该策略将每个动作分支一步,并对其对发出的词元的影响进行评分。这些收益教会策略何时记忆比进一步推理更有用,而通过反事实回忆的梯度则教会哪些中间状态应保留在记忆中供将来使用。在六个通用和长上下文推理基准上,1.4B的LatentHarness分别相对最强基线提高了2.8%和10.0%,并且运行速度比最强长上下文基线快5.9倍。
英文摘要
Long-context reasoning faces two complementary bottlenecks: retaining evidence across long inputs and sustaining computation across many reasoning steps. Existing approaches largely address them separately, with external memory extending access to distant evidence and latent reasoning compressing multi-step computation. We introduce LatentHarness, which unifies memory access and latent reasoning as sequential latent action selection. At each internal step, the model chooses THINK for further computation, RECALL from a fast-weight memory of input evidence and intermediate reasoning states, or EXIT to emit the next token. We train this policy with counterfactual policy distillation, which branches every action for one step and scores its effect on the emitted token. These gains teach the policy when memory is more useful than further reasoning, while gradients through counterfactual recall teach which intermediate states should be retained in memory for future use. Across six general and long-context reasoning benchmarks, LatentHarness at 1.4B improves on the strongest baselines by 2.8% and 10.0% relative, respectively, and runs 5.9x faster than the strongest long-context baseline.
CommentsWork in progress