散度几何:追踪隐状态轨迹以实现自适应多轮推理
Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning
- Lancaster University(兰卡斯特大学)
- University of Huddersfield(哈德斯菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究将多轮推理建模为LLM的隐状态轨迹,提出时间曲率与方差斜率两种几何信号,分解推理动作链,提升τ-Bench任务成功率并降低token成本。
AI中文摘要:
大语言模型(LLM)智能体需要在严格资源约束下,于长期多轮交互中维持与目标一致的推理。然而随着多轮上下文的累积,可能会破坏底层LLM对早期轮次任务相关信息的内部表征,模糊建设性推理与表征漂移的边界。我们将多轮推理建模为底层LLM的隐状态轨迹,该轨迹由两个互补信号表征:捕捉轮次间更新方向一致性的时间曲率,以及衡量探索空间扩张或收缩的方差斜率。在四项任务和三种底层LLM上,我们观察到这些几何信号能在推理完成前区分正确与错误的推理片段。我们进一步将每个片段分解为由四种动作(读取Read、写入Write、响应Respond、转移Transfer)构成的三动作链,且表明可分性依赖于动作,不同信号可区分各种链模式。实验表明,轨迹几何可识别推理过程中的关键轮次,将τ-Bench上的任务成功率从24.1%提升至39.6%,同时降低11.2%的token成本。
英文摘要:
LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant information from earlier turns, blurring the boundary between constructive reasoning and representation drift. We formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space. Across four tasks and three underlying LLMs, we observed that these geometric signals distinguish between correct and incorrect episodes prior to completion. We further decompose each episode into three-action chains formed from four actions (Read, Write, Respond, Transfer) and show that separability is action-dependent, with different signals distinguishing various chain patterns. Our experiments demonstrate that trajectory geometry can identify critical turns in the reasoning process, increasing task success rates on $τ$-Bench from 24.1% to 39.6% while reducing token cost by 11.2%.