发表机构
King Abdullah University of Science and Technology (KAUST); Edge Hill University(阿卜杜拉国王科技大学; 边山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出免训练的答案收敛停止规则,通过测量模型答案的自信与稳定状态实现自适应早停,在长上下文阅读中保持准确率并节省计算,无需额外训练。
AI 中文摘要
语言模型通常按块顺序处理长输入,但在获得足够证据后继续阅读会浪费计算资源。现有的停止机制要么从内部激活中学习充分性,要么训练一个退出门控,而一个更简单的替代方案是询问模型是否已读够。我们引入了答案收敛停止(ACS),一种免训练的停止规则,它通过测量而非询问来判断。在每个块之后,它探测冻结模型的当前答案状态,并在该状态既自信又稳定时停止。该规则仅需要输出端的生成和令牌对数概率,没有训练组件,并且在模型和基准测试中使用一个共享配置。由于停止策略可以通过过早停止来节省计算,我们在有证据位置的情况下评估停止决策本身。在完整的LongBench-v2上使用两个前沿模型,ACS是唯一达到或超过全读准确率的停止策略。此外,在250个S-NIAH问题中,来自两个家族的五个模型的ACS过早停止率范围为0%至12%,而口头门控的过早停止率为8.4%至45.6%。综合来看,ACS表明,通过适当利用冻结模型的输出信号,我们可以在无需额外训练的情况下实现自适应停止等有利行为。
英文摘要
Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model whether it has read enough. We introduce Answer-Convergence Stopping (ACS), a training-free stopping rule that measures rather than asks. After each chunk, it probes the frozen model's current answer state and stops when that state is both confident and stable. The rule requires only output-side generation and token log probabilities, has no trained components, and uses one shared configuration across models and benchmarks. Because a stopping policy can save computation simply by stopping too early, we evaluate the stopping decision itself using evidence position where available. On the full LongBench-v2 with two frontier models, ACS is the only stopping policy that matches or exceeds full-reading accuracy. Furthermore, across 250 S-NIAH questions, the premature stopping rate for ACS across five models from two families ranges from 0% to 12%, compared to 8.4% to 45.6% for the verbalized gate. Taken together, ACS reveals that by properly utilizing the output signals of frozen models, we can achieve favorable behaviors like adaptive stopping without the need for additional training.