arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BODHI:大型语言模型是否会分支拓展并发现异构推理?

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

Soumadeep Saha, Krish Sharma, Akshay Chaturvedi, Nicholas Asher

arXiv 2608.02867首次发表:更新:

发表机构

ANITI, Université de Toulouse; LINAGORA Labs(图卢兹大学ANITI; LINAGORA实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过迷宫求解实验和BODHI-Trees分析,发现RLVR训练的LLMs存在策略熵崩溃,其采样效率提升以推理分支多样性下降为代价。

AI 中文摘要

尽管带可验证奖励的强化学习(RLVR)已提升了大型语言模型(LLMs)在多种推理任务上的性能,但关于RLVR是拓展了推理能力边界,还是仅提升了采样效率,仍存在重大争议。本文中,我们通过受控迷宫求解实验,并基于语义等价从数学推理轨迹中提取树结构(BODHI-Trees),来研究RLVR训练的LLMs在测试时探索的本质。这有助于我们区分风格变化产生的熵与真正的推理分支。我们的发现表明,RLVR模型中观察到的策略熵崩溃并非仅为句法层面,还伴随语义分支熵的显著降低。RLVR虽提升了对环境约束的遵守和回溯能力,但收缩了延续空间;我们提供的证据表明,这可能是RLVR采样效率提升的原因,尽管以真正的 rollout 多样性为代价。

英文摘要

Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning capability boundary, or just improves sampling efficiency. In this paper, we investigate the nature of test-time exploration in RLVR-trained LLMs by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence. This helps us delineate between entropy arising from stylistic variations and genuine inferential branching. Our findings demonstrate that the policy entropy collapse observed in RLVR models is not merely syntactic, and is accompanied by a significant reduction in semantic branching entropy. While RLVR improves adherence to environmental constraints and backtracking capabilities, it constricts the space of continuations; we provide evidence suggesting that this might be responsible for the sample efficiency gains of RLVR, albeit at the cost of genuine rollout diversity.

Comments16 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑