发表机构
UC San Diego; Adobe Research; University of New South Wales(加州大学圣地亚哥分校; Adobe研究院; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型推理过程不可控问题,提出SOPHIA方法,通过将推理轨迹视为潜在状态序列,构建引导向量库,在推理时进行干预,可检测并防止自我循环,提升推理质量。
AI 中文摘要
扩展推理已成为前沿大语言模型(LLMs)的标准,但模型生成的轨迹在很大程度上仍无法控制。现有塑造模型推理方式的方法是基于提示的,在输入层面操作,无法对推理过程本身进行细粒度控制。相关工作分析并发现了大语言模型推理轨迹中的潜在转换动态。在此基础上,我们对这些状态进行统计表征,发现失败轨迹会陷入自我循环。为干预这些失败,我们提出了SOPHIA:通过隐藏状态干预和激活来引导推理过程。我们将每个推理轨迹视为潜在状态序列,分类前缀到潜在状态,记录步骤级转换,构建引导向量库。在推理时,控制器推断当前状态,给定目标状态检索相应向量,还能从转换结构中在线检测自我循环。实验表明,我们的方法能可靠干预自我循环失败,引导向量可推广到不同状态对,细粒度可控性带来更好的推理质量。
英文摘要
Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable. Existing methods for shaping how a model reasons are prompt based approaches and operate at the input level, offering no fine-grained control over the reasoning process itself. Related work analyzes and discovers latent transition dynamics in the reasoning traces from Large Language Models. Building on this, we statistically characterize these states, and show that failure trajectories get stuck in self-loops, exhausting the token budget without progress toward the final answer. To intervene on these failures, We propose SOPHIA: Steering Of reasoning Processes via Hidden-state Intervention and Activations. We treat each reasoning trace as a sequence of latent states rather than an unstructured texts, and investigate whether inference time interventions can provide fine-grained control over the self-looping reasoning process. We classify every prefix to a latent state, record step level transitions, and use them to construct a bank of steering vectors indexed by state pairs. At inference time, a controller infers the current state and, given a target state, retrieves the corresponding vector and can also detect self-loops online from the transition structure to prevent the model from sinking into a reasoning black hole. Through extensive experiments, our method reliably intervenes on self-loop failures, with steering vectors that generalize to different state pairs. End task accuracy and token efficiency indicate that fine-grained controllability results in better reasoning quality.