arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于答案探测引导的大语言模型多样化解决方案探索搜索

Answer Probing-Guided Search for Diverse Solution Exploration of LLMs

Yi Fang, Que Shen, Chengpeng Li, Boyi Deng, Wei Shi, Wenjie Wang, Fuli Feng, Fengli Xu, Dayiheng Liu

arXiv 2608.30345首次发表:更新:

发表机构

University of Science and Technology of China; Zhongguancun Academy; Alibaba Group; Shanghai Jiao Tong University; Tsinghua University(中国科学技术大学; 中关村学院; 阿里巴巴集团; 上海交通大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLMs推理易收敛于单一解的问题,提出Answer Probing引导的树搜索APTS,经实验证实可提升多推理任务的解决方案多样性,具备有效性与鲁棒性。

AI 中文摘要

生成多个多样化且高质量的解决方案对代码测试生成、药物发现等诸多应用具有重要价值。然而,大语言模型(LLMs)在推理过程中倾向于收敛到单一高置信度解决方案,限制了对其他有效解决方案路径的探索。现有的测试时方法通过类树搜索并使用响应级语义嵌入修剪语义相似分支来促进多样性,但我们发现此类嵌入易受语言和风格相似性干扰,难以区分真正不同的解决方案路径。为解决该问题,我们引入答案探测(Answer Probing),即探测LLMs从中间推理路径可能得出的潜在答案。我们证明,探测答案的隐藏状态比语义嵌入更能有效区分不同解决方案路径,且探测答案的困惑度可作为推理正确性的实用代理。基于这些发现,我们提出答案探测引导的树搜索(APTS),通过探测答案的隐藏状态相似性和困惑度来引导树搜索。在两个LLMs的三个推理任务上的实验表明,APTS持续提升了解决方案多样性,证明了其有效性和鲁棒性。

英文摘要

Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery. However, Large Language Models (LLMs) tend to converge on a single high-confidence solution during inference, limiting exploration of alternative valid solution paths. Existing test-time methods promote diversity through tree-like search and prune semantically similar branches using response-level semantic embeddings. However, we find that such embeddings are easily confounded by linguistic and stylistic similarities, making it difficult to distinguish genuinely distinct solution paths. To address this, we introduce Answer Probing, which probes the potential answer an LLM would reach from an intermediate reasoning path. We demonstrate that the hidden states of probed answers more effectively differentiate distinct solution paths than semantic embeddings, and the perplexity of probed answers serves as a practical proxy for reasoning correctness. Based on these findings, we propose Answer Probing-Guided Tree Search (APTS), which guides the tree search by the probed answers' hidden state similarity and perplexity. Experiments on three reasoning tasks across two LLMs show that APTS consistently enhances solution diversity, demonstrating its effectiveness and robustness.

CommentsAccepted to the EMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑