arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

选择性再生解码:推理时推理的轨迹级干预

Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning

Sophia Xiao Pu, Yumo Xu, Sailik Sengupta, Millennium Bismay, Ruixue Lian, James Gung, Yi-an Lai, Arshit Gupta

arXiv 2608.24338首次发表:更新:

发表机构

University of California, Santa Barbara; Netflix; Amazon Science; Meta(加利福尼亚大学圣巴巴拉分校; 网飞公司; 亚马逊科学研究院; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出选择性再生解码(SRD),通过对候选轨迹做片段级干预,在低计算量下以更少生成代币达到Best-of-N准确率,提升样本效率,开辟了推理时推理的准确率-计算权衡新领域。

AI 中文摘要

推理时解码方法通过探索多个候选轨迹提升大语言模型(LLM)的推理能力,但将每个轨迹视为原子单元:要么完整保留,要么不可逆地丢弃。这会浪费计算资源在部分有潜力的候选上,这些候选的高质量前缀会随退化后缀一同被丢弃。我们提出选择性再生解码(SRD),其对每个候选进行分类:丢弃、保留,或仅对边界候选的退化后缀部分进行优化,同时保留有用前缀,且无需更大的目标模型。在温和假设下,SRD相较于拒绝采样可实现可证明的1.28至1.36倍的样本效率提升,且预期轨迹质量严格更高,该提升随候选池规模增大而增长。在MATH500、GPQA Diamond、HotpotQA及AlpacaEval上,搭配多个生成奖励模型对时,SRD达到了Best-of-N的准确率,但生成的代币量显著更少,且在低计算 regimes下优于推测拒绝。通过实现片段级干预而非全轨迹选择,SRD为推理时推理的准确率-计算权衡开辟了此前未充分探索的领域。

英文摘要

Inference-time decoding methods improve LLM reasoning by exploring multiple candidate trajectories, yet treat each trajectory as atomic: either retaining it whole or discarding it irreversibly. This wastes computation on partially promising candidates whose high-quality prefixes are abandoned alongside degraded suffixes. We introduce Selective Regenerative Decoding (SRD), which routes each candidate to discard, keep, or refine only the degraded portion of the suffix while preserving the useful prefix of borderline candidates, without requiring a larger target model. Under mild assumptions, SRD achieves a provable 1.28-to-1.36-fold gain in sample efficiency over rejection sampling with strictly higher expected trajectory quality, with the gain growing as the candidate pool grows. Across MATH500, GPQA Diamond, HotpotQA, and AlpacaEval with multiple generation-reward model pairs, SRD matches Best-of-N accuracy with substantially fewer generated tokens and outperforms speculative rejection in low-compute regimes. By enabling segment-level intervention rather than whole-trajectory selection, SRD opens a previously underexplored region of the accuracy-compute tradeoff for inference-time reasoning.

CommentsLarge Language Model, Test-Time Decoding Technique, Reasoning, 20 pages; Submitted to ARR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑