发表机构
Beijing Institute for General Artificial Intelligence; Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, Beijing University of Posts and Telecommunications; School of Artificial Intelligence for Science, Peking University(北京通用人工智能研究院; 中国科学院自动化研究所; 北京邮电大学人工智能学院; 北京大学科学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GradCuit在测试时插入可优化潜在状态,实现序列级信用分配,在多基准上准确率优于同类方法,且鲁棒性、可解释性更强,为LLM测试时推理扩展提供新方向。
AI 中文摘要
基于优化的潜在推理在测试时通过优化实例特定的连续状态来改进大语言模型(LLM)的输出,同时保持模型参数冻结。然而,现有方法通常将这些状态通过解码后的标记与推理轨迹关联,导致序列级信用分配间接,且模糊了潜在更新如何塑造后续推理。我们提出GradCuit(梯度通过电路),它在提示的隐藏表示和生成的延续之间的选定Transformer层插入可优化的潜在状态。因果自注意力为每个延续标记的对数概率提供了一条通过剩余Transformer块连接到所有先前潜在状态的可微路径,使得来自整个延续的奖励加权梯度能够直接分配给潜在变量。在五个指令微调的骨干模型、三个推理基准和两种答案格式上,GradCuit达到了64.5%的平均准确率,比思维链(Chain-of-Thought)提示高出6.6个百分点,比最强的竞争方法高出2.4个百分点。GradCuit还展现出更强的鲁棒性:在七个学习率设置下,它始终优于LatentSeek,同时将准确率的标准差从1.53降至0.82,甚至其随机游走变体仍与LatentSeek具有竞争力。在可解释性方面,标记级梯度归因显示潜在影响集中在推理连接标记上,而层分析确定Transformer的中低层是最有效的优化空间。通过直接从结果反馈中优化内部推理,GradCuit开辟了鲁棒且可解释的测试时扩展的新维度,其中LLM调整自身推理方式,而非仅重新生成、采样或重新排序输出。
英文摘要
Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these states to the reasoning trajectory through decoded tokens, making sequence-level credit assignment indirect and obscuring how latent updates shape subsequent reasoning. We introduce GradCuit (gradient through circuit), which inserts optimizable latent states at a selected Transformer layer between the hidden representations of the prompt and the generated continuation. Causal self-attention provides every continuation-token log-probability with a differentiable path to every preceding latent state through the remaining Transformer blocks, enabling reward-weighted gradients from the entire continuation to be assigned directly to the latents. Across five instruction-tuned backbones, three reasoning benchmarks, and two answer formats, GradCuit achieves an average accuracy of 64.5%, outperforming chain-of-thought prompting by 6.6 percentage points and the strongest competing method by 2.4 points. GradCuit also demonstrates greater robustness: across seven learning-rate settings, it consistently outperforms LatentSeek while reducing the standard deviation of accuracy from 1.53 to 0.82, and even its random-walk variant remains competitive with LatentSeek. For interpretability, token-level gradient attribution reveals that latent influence concentrates on reasoning-connector tokens, while layer analysis identifies early-to-middle Transformer layers as the most effective optimization space. By directly optimizing internal reasoning from outcome feedback, GradCuit opens a new axis of robust and interpretable test-time scaling, where LLMs adapt how they reason rather than merely regenerate, sample, or rerank outputs.