A*-Thought-V2:通过LLM的几何动力学实现高效潜在推理
A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
- Beijing University of Posts and Telecommunications(北京邮电大学)
- The Hong Kong Polytechnic University(香港理工大学)
- Tsinghua University(清华大学)
- JIUTIAN Research(中移九天)
- OpenBMB
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
A*-Thought-V2通过几何动力学引导的显式-隐式交错潜在架构,将偏离方向的推理步骤压缩为潜在令牌,在提升准确率的同时大幅降低计算成本。
AI中文摘要:
思维链(CoT)提升了大语言模型(LLMs)的推理能力,但会带来大量的计算和上下文成本。现有方法要么通过硬剪枝丢失中间信息,要么缺乏连续压缩的原则性标准。我们提出了A*-Thought-V2,一种由LLM几何动力学引导的框架,将CoT建模为隐藏状态轨迹,并用显式-隐式交错潜在架构取代硬删除。在将问题、步骤和解决方案表示投影到3D PCA空间后,它衡量每个局部转换与全局问题到解决方案方向之间的对齐程度。对齐的步骤保持显式文本,而偏离的步骤被压缩为连续潜在令牌。方向角度同时捕获局部语义和推理动态:小角度表示直接执行和答案形成,而大角度更频繁地涉及检查、纠正和分支探索;它们的时间变化揭示了探索、收敛和细化阶段。为了训练该架构,我们引入了逐步嵌入强制,将每个冗余步骤汇聚为单个潜在嵌入,以及标签强制,用软多模态词汇分布而非硬独热标签来监督该潜在令牌。在Qwen3.5-9B和Qwen3.6-27B上跨六个域内和域外基准的实验表明,A*-Thought-V2将平均准确率提高多达2.6%,同时将响应长度减少多达一半,将每计算单元准确率提高2.29倍,并将预处理和训练时间分别减少94.6%和最多80.3%。表示分析表明,潜在状态形成一个与文本状态不同的紧凑区域,而潜在令牌位置处的更高熵反映了更广泛的软目标,鼓励更丰富的步骤级特征学习。
英文摘要:
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture. After projecting question, step, and solution representations into a 3D PCA space, it measures alignment between each local transition and global question-to-solution direction. Aligned steps remain explicit text, whereas deviating steps are compressed into continuous latent tokens. Directional angles capture both local semantics and reasoning dynamics: small angles indicate direct execution and answer formation, while large angles more frequently involve checking, correction, and branch exploration; their temporal variation reveals exploration, convergence, and refinement stages. To train this architecture, we introduce stepwise embedding forcing, which pools each redundant step into a single latent embedding, and label forcing, which supervises that latent token with a soft multi-modal vocabulary distribution instead of a hard one-hot label. Experiments on Qwen3.5-9B and Qwen3.6-27B across six in-domain and out-of-domain benchmarks show that A*-Thought-V2 improves average accuracy by up to 2.6% while reducing response length by up to half, increasing Accuracy per Computation Unit by 2.29$\times$, and reducing preprocessing and training time by 94.6% and up to 80.3%, respectively. Representation analyses suggest that latent states form a compact region distinct from textual states, while higher entropy at latent-token positions reflects broader soft targets that encourage richer step-level feature learning.