arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReSolve:通过选择性生成式调节重用候选推理

ReSolve: Reusing Candidate Reasoning through Selective Generative Moderation

Bangji Yang, Jiajun Fan, Hongbo Ma, Xi Zhu, Weizhi Zhang, Minghao Guo, Ye Li, Hamid Palangi, Jiaxuan You

arXiv 2610.01140首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; Tsinghua University; University of Illinois Chicago; Rutgers University; Google(伊利诺伊大学厄巴纳-香槟分校; 清华大学; 伊利诺伊大学芝加哥分校; 罗格斯大学; 谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReSolve通过选择性生成式调节重用候选推理,在竞赛数学问题上以更少令牌取得优于投票的准确率,证明候选推理是可重用的计算资源。

AI 中文摘要

对多个解进行采样会在中间推导、未完成的论证以及最终答案上耗费计算资源。我们提出ReSolve,一种免训练的推理过程,通过选择性生成式调节来重用这些候选推理。当候选答案不一致或缺乏可解析的答案时,一个答案分布控制器调用模型检查现有推导,并将生成的解纳入有界循环。在130道竞赛数学问题上的混合评分中,使用两个独立采样的候选池进行评估,ReSolve分别获得100和99个正确答案,而相同四个候选的投票分别获得91和92个正确答案,且两个池中均未出现相对于投票的从正确到错误的转变。八样本自洽性获得94和96个正确答案,但消耗了显著更多的令牌;ReSolve在两次评估中分别使用46.3%和47.2%更少的令牌。一项受控消融移除了可见推导,同时保留答案键、投票计数和每状态输出上限规则,尽管计算量增加,准确率从100降至93个正确答案。选择性和始终开启的均匀调节均解决97个问题,而选择性将调节令牌减少约54%,总管道令牌减少6.2%。这些结果支持候选推理作为可重用的推理计算。它们并未建立相对于额外采样的准确性优势或专门路线指令的独特益处。

英文摘要

Sampling multiple solutions spends computation on intermediate deductions and unfinished arguments as well as final answers. We introduce ReSolve, a training-free inference procedure that reuses this candidate reasoning through selective generative moderation. An answer-distribution controller invokes a model to examine existing derivations when candidates disagree or lack a parseable answer, then incorporates the generated solution into a bounded loop. Under Hybrid scoring on 130 competition-mathematics problems evaluated with two independently sampled candidate pools, ReSolve obtains 100 and 99 correct answers, compared with 91 and 92 for voting over the same four candidates, with no correct-to-incorrect changes relative to that vote in either pool. Eight-sample self-consistency obtains 94 and 96 correct answers while consuming substantially more tokens; ReSolve uses 46.3% and 47.2% fewer tokens in the two evaluations. A controlled ablation removes visible derivations while retaining answer keys, vote counts, and the per-state output-cap rule, reducing accuracy from 100 to 93 correct despite increasing computation. Selective and always-on Uniform moderation both solve 97 problems, while selectivity reduces moderation tokens by approximately 54% and total pipeline tokens by 6.2%. These results support candidate reasoning as reusable inference computation. They do not establish an accuracy advantage over additional sampling or a distinct benefit from specialized route instructions.

CommentsCorrected a typo in an author's name. No changes to the paper content

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑