arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07821cs.CLcs.AIcs.LG

A*-Thought-V2:通过LLM的几何动力学实现高效潜在推理

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • The Hong Kong Polytechnic University(香港理工大学)
  • Tsinghua University(清华大学)
  • JIUTIAN Research(中移九天)
  • OpenBMB

机构由 AI 辅助整理,请以论文原文为准。

Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liu… 展开作者

Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He

AI总结:

A*-Thought-V2通过几何动力学引导的显式-隐式交错潜在架构,将偏离方向的推理步骤压缩为潜在令牌,在提升准确率的同时大幅降低计算成本。

AI中文摘要:

思维链(CoT)提升了大语言模型(LLMs)的推理能力,但会带来大量的计算和上下文成本。现有方法要么通过硬剪枝丢失中间信息,要么缺乏连续压缩的原则性标准。我们提出了A*-Thought-V2,一种由LLM几何动力学引导的框架,将CoT建模为隐藏状态轨迹,并用显式-隐式交错潜在架构取代硬删除。在将问题、步骤和解决方案表示投影到3D PCA空间后,它衡量每个局部转换与全局问题到解决方案方向之间的对齐程度。对齐的步骤保持显式文本,而偏离的步骤被压缩为连续潜在令牌。方向角度同时捕获局部语义和推理动态:小角度表示直接执行和答案形成,而大角度更频繁地涉及检查、纠正和分支探索;它们的时间变化揭示了探索、收敛和细化阶段。为了训练该架构,我们引入了逐步嵌入强制,将每个冗余步骤汇聚为单个潜在嵌入,以及标签强制,用软多模态词汇分布而非硬独热标签来监督该潜在令牌。在Qwen3.5-9B和Qwen3.6-27B上跨六个域内和域外基准的实验表明,A*-Thought-V2将平均准确率提高多达2.6%,同时将响应长度减少多达一半,将每计算单元准确率提高2.29倍,并将预处理和训练时间分别减少94.6%和最多80.3%。表示分析表明,潜在状态形成一个与文本状态不同的紧凑区域,而潜在令牌位置处的更高熵反映了更广泛的软目标,鼓励更丰富的步骤级特征学习。

英文摘要:

Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture. After projecting question, step, and solution representations into a 3D PCA space, it measures alignment between each local transition and global question-to-solution direction. Aligned steps remain explicit text, whereas deviating steps are compressed into continuous latent tokens. Directional angles capture both local semantics and reasoning dynamics: small angles indicate direct execution and answer formation, while large angles more frequently involve checking, correction, and branch exploration; their temporal variation reveals exploration, convergence, and refinement stages. To train this architecture, we introduce stepwise embedding forcing, which pools each redundant step into a single latent embedding, and label forcing, which supervises that latent token with a soft multi-modal vocabulary distribution instead of a hard one-hot label. Experiments on Qwen3.5-9B and Qwen3.6-27B across six in-domain and out-of-domain benchmarks show that A*-Thought-V2 improves average accuracy by up to 2.6% while reducing response length by up to half, increasing Accuracy per Computation Unit by 2.29$\times$, and reducing preprocessing and training time by 94.6% and up to 80.3%, respectively. Representation analyses suggest that latent states form a compact region distinct from textual states, while higher entropy at latent-token positions reflects broader soft targets that encourage richer step-level feature learning.

补充信息

↑