arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

束搜索作为基于反事实上下文的测试时自蒸馏

Beam Search as Test-Time Self-Distillation via Counterfactual Contexts

Su Ee Tan, Xiaotong Ji, Rasul Tutunov, Haitham Bou-Ammar, Matthieu Zimmer

arXiv 2609.37041首次发表:更新:

发表机构

UCL Centre for AI(伦敦大学学院人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出测试时自蒸馏方法,利用反事实上下文定义奖励信号,通过束搜索近似全局重加权,在推理时无需训练即可提升数学、代码和科学问答性能。

AI 中文摘要

自蒸馏微调(SDFT)使语言模型能够充当自己的教师:通过以演示为条件,模型通过逐点互信息产生隐式奖励,从而在无需外部监督的情况下指导在线学习。然而,SDFT在训练时运行:它需要梯度更新和专家演示的访问权限,使其在推理时无法应用。我们提出测试时自蒸馏,一种解码时方法,从自蒸馏框架中提取引导信号,无需任何参数更新、奖励模型或训练数据。我们的关键见解是,反事实上下文,即假设性地将模型预设为优秀与较差推理的固定文本模板,可以替代演示。候选答案在这两种反事实条件下的对数几率比定义了一种新的奖励信号。我们推导了在此奖励下的最优KL正则化策略,其形式为对基础分布的吉布斯重新加权。至关重要的是,这种重新加权是全局性的:它不能分解为独立的逐词元操作而不忽略未来轨迹质量。因此,我们通过束搜索来近似目标分布。在数学推理(MATH500)、代码生成(HumanEval)和研究生级科学问答(GPQA)上的实验,跨越多个模型规模,表明测试时自蒸馏在平均性能上优于标准采样、低温采样、束搜索和幂采样基线,证明了自蒸馏原理可以在推理时得以实现。

英文摘要

Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient updates and access to expert demonstrations, making it inapplicable at inference. We propose test-time self-distillation, a decoding-time method that extracts a steering signal from the self-distillation framework without any parameter updates, reward models, or training data. Our key insight is that counterfactual contexts, i.e. fixed textual templates that hypothetically prime the model for excellent versus poor reasoning, can substitute for the demonstration. The log-odds ratio of a candidate answer under these two counterfactual conditions defines a new reward signal. We derive the optimal KL-regularized policy under this reward, which takes the form of a Gibbs reweighting of the base distribution. Crucially, this reweighting is global: it cannot be decomposed into independent per-token operations without ignoring future trajectory quality. We therefore approximate the target distribution via beam search. Experiments on mathematical reasoning (MATH500), code generation (HumanEval), and graduate-level science QA (GPQA) across multiple model scales show that test-time self-distillation improves over standard sampling, low temperature, beam search and power sampling baselines on average, demonstrating that the self-distillation principle can be operationalized at inference time.

CommentsAccepted at NeurIPS 2026 Workshop on Towards Test-Time Continual Learning Agents

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑