AI 中文总结
该研究针对芯片设计中高覆盖率测试平台激励生成任务,提出CHORUS后训练框架,整合分阶段SFT与密集奖励RL得到的互补专家,构建4B模型,在CVDP-ECov上的Pass@1指标优于671B参数的DeepSeek-R1。
AI 中文摘要
大型语言模型(LLM)已推动代码生成技术发展,其中可执行反馈相比单纯文本模仿能提供更可靠的学习信号。硬件验证是代码生成的重要应用,占现代芯片设计工作的很大一部分,而高覆盖率测试平台激励生成是其关键任务。我们提出CHORUS,这是一种后训练框架,其性能超越了传统的监督微调(SFT)到强化学习(RL)的流水线所能达到的水平。CHORUS基于两个观察结果:第一,分阶段SFT会产生行为多样的检查点,而密集奖励RL将它们转化为强大的专家,这些专家具有相当的综合性能,但在任务层面的优势各不相同;第二,这些互补优势可以通过无需训练的模型合并或进一步后训练来利用,从而优于最佳的单个专家。通过将生成的专家整合为一个4B参数的单一模型,CHORUS在CVDP-ECov上实现了88.0%的Pass@1指标,比DeepSeek-R1(671B)高出13.5个百分点。
英文摘要
Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and accounts for a substantial fraction of modern chip design effort, with high-coverage testbench stimulus generation as a key task. We present CHORUS, a post-training framework that pushes performance beyond what a conventional supervised fine-tuning (SFT)-to-reinforcement learning (RL) pipeline achieves. CHORUS builds on two observations. First, staged SFT produces behaviorally diverse checkpoints, and dense-reward RL turns them into strong experts with comparable aggregate performance but distinct task-level strengths. Second, these complementary strengths can be exploited through either training-free model merging or further post-training to outperform the best individual expert. By consolidating the resulting specialists into a single 4B model, CHORUS achieves 88.0% Pass@1 on CVDP-ECov, outperforming DeepSeek-R1 (671B) by 13.5 percentage points.