发表机构
Virginia Tech; University of Wisconsin Madison; Dartmouth College(弗吉尼亚理工大学; 威斯康星大学麦迪逊分校; 达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SAGE通过代数稀疏化和双曲结构引导,统一缓解长程推理中的探索与复合偏差,在12个基准上超越基线,并在Andrews-Curtis问题上实现8倍改进。
AI 中文摘要
在稀疏奖励机制下,长程推理仍然是大型语言模型(LLMs)面临的核心挑战。我们认为,这种脆弱性源于复杂推理空间所引发的两种偏差:探索偏差,即模型倾向于选择局部看似合理但结构上不稳定的分支;以及复合偏差,即微小的局部偏差在深度上累积并抑制了稀有奖励的出现。我们引入符号闭包分析(SCA)作为理论视角,用以刻画分支结构和稀疏奖励如何在具有局部可容许性的长程推理中引发这些偏差,并作为在形式化程度较低的推理任务中设计结构先验的原则。基于此分析,我们提出SAGE(结构可容许性引导探索),一个统一框架,通过注入结构引导来缓解长程推理中的探索偏差和复合偏差。SAGE结合了两种互补的结构引导:代数稀疏化,将局部可容许候选投影到算子索引的代数子空间以抑制虚假分支并缓解探索偏差;以及双曲结构引导,将推理状态嵌入负曲率空间以提供密集的深度方向信号并缓解复合偏差。在12个基准测试和7个模型家族中,SAGE优于竞争性基线。特别是,在Andrews-Curtis问题(一个开放的真实世界长程任务)上,SAGE实现了高达8倍的改进。代码可在以下网址获取:此https URL。
英文摘要
Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compounding bias, where small local deviations accumulate across depth and suppress rare rewards. We introduce Symbolic Closure Analysis (SCA) as a theoretical lens characterizing how branching structures and sparse rewards induce these biases in long-horizon reasoning with local admissibility, and as a design principle for structural priors in less formal reasoning tasks. Motivated by this analysis, we propose SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning. SAGE combines two complementary structural guidance: algebraic sparsification, which projects locally admissible candidates onto operator-indexed algebraic subspaces to suppress spurious branching and mitigate exploration bias, and hyperbolic structural guidance, which embeds reasoning states into a negatively curved space to provide dense depth-wise signals and mitigate compounding bias. Across 12 benchmarks and 7 model families, SAGE outperforms competitive baselines. In particular, SAGE achieves up to an 8-fold improvement on the Andrews-Curtis problem, an open real-world long-horizon task. Code is available at: https://github.com/Susan571/SAGE-NeurIPS2026.
CommentsAccepted by NeurIPS 2026