arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26457cs.LG

DHRCL:使用密集分层奖励与课程学习训练代码大语言模型

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning

Shuhang Wang, Ziming Li, Hui Cheng

AI总结:

该研究提出DHRCL框架,通过分层奖励与自动阶段课程学习训练代码大语言模型,经多模型规模实验验证其优势稳定,优于多种基线方法。

AI中文摘要:

强化学习是面向代码的大语言模型的自然后训练范式,因为生成的程序可通过解析、执行、单元测试和结构this http URL评估。现有方法通常依赖稀疏结果奖励或静态组合异构密集信号,尽管语法有效性、可执行性、功能正确性和结构组织描述了不同且逐步依赖的编程能力。我们提出DHRCL,一个具有密集分层奖励和课程学习的强化学习框架。DHRCL将反馈分解为语法验证、执行成功、单元测试通过率和基于AST的结构相似度,并通过三阶段语法、执行、通过与结构课程组织这些信号。阶段持续时间由最近的验证趋势自动确定,而非手动指定的能力阈值。我们进一步引入阶段感知的基于概率的令牌信用再分配机制,该机制遵循巩固到细化原则:在面向语法的优化中强调已建立的令牌模式,对非本地执行反馈应用均匀传播,并在最终功能优化中为较不确定的令牌决策分配更多信用或责备。在统一的Qwen3-8B和KodCode协议下,实验将DHRCL与二元、通过率、基于奖励模型和可验证密集奖励的基线进行比较。我们还在Qwen3-4B、Qwen3-8B和Qwen3-14B主干上评估DHRCL,表明其优势随模型容量增加保持一致。

英文摘要:

Reinforcement learning is a natural post-training paradigm for code-oriented large language models because generated programs can be evaluated through parsing, execution, unit tests, and structural analysis. However, existing methods often rely on sparse outcome rewards or statically combine heterogeneous dense signals, even though syntax validity, executability, functional correctness, and structural organization describe different and progressively dependent programming capabilities. We propose DHRCL, a reinforcement learning framework with Dense Hierarchical Rewards and Curriculum Learning. DHRCL decomposes feedback into syntax validation, execution success, unit-test pass rate, and AST-based structural similarity, and organizes these signals through a three-stage Syntax, Execution, Pass & Structural curriculum. Stage duration is determined automatically from recent validation trends rather than manually specified capability thresholds. We further introduce stage-aware probability-based token credit redistribution. The mechanism follows a consolidation-to-refinement principle: it emphasizes established token patterns during syntax-oriented optimization, applies uniform propagation for non-local execution feedback, and allocates more credit or blame to less-established token decisions during final functional optimization. Under a unified Qwen3-8B and KodCode protocol, the experiments compare DHRCL with binary, pass-rate, reward-model-based, and verifiable dense-reward baselines. We further evaluate DHRCL across Qwen3-4B, Qwen3-8B, and Qwen3-14B backbones, showing that its advantage remains consistent as model capacity increases.

↑