arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CORE:面向可验证语言模型搜索的冲突导向推理消除

CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search

Siyu Song, Rui Xu, Jia Lin, Kai Liu, Weifang Wang

arXiv 2609.39069首次发表:更新:

发表机构

Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CORE通过验证器认证冲突核心并回跳至相关决策,避免重复错误,在完备搜索下减少验证调用并提升推理任务成功率。

AI 中文摘要

测试时推理系统在遇到失败时,常常通过重启或修订最近一步来应对,即使错误是由更早的决策导致的。我们提出CORE,一种搜索控制器,它向验证器请求一个经过认证的冲突核心,回跳到该核心中的最新决策,并缓存该冲突以避免重复出现。在可靠的验证、有限的分支和深度以及穷举提议的条件下,无上限的搜索是完备的,且永远不会剪除有效解。在2000个带有匹配提议和精确验证器的植入式图着色实例上,与按时间顺序修复相比,CORE在30个变量时减少了39.8%的中位验证器调用次数,在36个变量时减少了35.0%;缓存进一步优于仅回跳。在五个推理任务中,CORE在使用Qwen2.5-7B-Instruct时达到75.9%的平均成功率,在使用Qwen3-8B时达到84.2%,而思维树分别为72.5%和81.8%。它在两个骨干模型上也使用了更少的验证器调用和生成的令牌。这些结果表明,使用经过认证的失败解释来引导语言模型搜索是有价值的。

英文摘要

Test-time reasoning systems often respond to failure by restarting or revising the latest step, even when an earlier decision caused the error. We introduce CORE, a search controller that requests a certified conflict core from a verifier, backjumps to the latest decision in that core, and caches the conflict to avoid repeating it. Under sound verification, finite branching and depth, and exhaustive proposals, the uncapped search is complete and never prunes a valid solution. On 2,000 planted graph-coloring instances with matched proposals and an exact verifier, CORE reduces median verifier calls by 39.8% at 30 variables and 35.0% at 36 variables relative to chronological repair; caching further improves on backjumping alone. Across five reasoning tasks, CORE achieves 75.9% mean success with Qwen2.5-7B-Instruct and 84.2% with Qwen3-8B, compared with 72.5% and 81.8% for Tree of Thoughts. It also uses fewer verifier calls and generated tokens on both backbones. These results show the value of using certified failure explanations to direct language-model search.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑