面向Verilog代码生成的执行锚定幻觉校准重排序
Execution-Anchored Hallucination Calibration Reranking for Verilog Code Generation
浏览论文内容
中文总结 AI 辅助
该研究针对LLM生成Verilog代码时的性能问题,提出EAHC重排序框架,通过双通道架构融合执行与推理信号,解决现有方法的域迁移差和推理幻觉问题。
中文摘要 AI 辅助
大语言模型(LLMs)在代码生成领域展现出卓越能力,但在Verilog等低资源硬件描述语言上的性能显著下降。尽管多候选采样提高了生成正确解的可能性,但自动选择最优候选仍是未解决的挑战。通过对9个模型和2个基准进行系统实证研究,我们发现两个关键局限:(1)现有基于执行的重排序方法依赖测试平台的通过/失败结果,因生成的测试平台质量低而域迁移能力差;(2)作为评判者的LLM存在推理幻觉,对执行等价的代码会产生不一致的判断。这些发现揭示了两类带有正交误差的信号:执行信号(确定性但测试平台覆盖率有限)和推理信号(语义丰富但易产生幻觉)。它们的正交性表明应结合这两类信号,但实验中让推理器直接观察执行结果只会将判断锚定在测试结果上;因此我们独立获取两类信号,仅在决策阶段融合。基于这些见解,我们提出EAHC(Execution-Anchored Hallucination Calibration)重排序框架,将推理判断锚定到执行行为,使执行等价的候选获得一致分数,该框架实现了双通道架构:EAHC-R(一个4B参数的推理判别器)和EAHC-T(利用RAG进行执行验证的测试平台生成器)。
英文摘要
Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation, yet their performance degrades significantly on low-resource Hardware Description Languages such as Verilog. While multi-candidate sampling improves the likelihood of generating correct solutions, au-tomatically selecting the optimal candidate remains an open challenge. Through a systematic empirical study across nine models and two benchmarks, we identify two critical limitations:(1) existing execution-based reranking methods, which rely on testbench pass/fail outcomes, exhibit poor domain transferability due to low-quality generated testbenches; and (2) LLM-as-a-Judge suffers from reasoning hallucination, producing incon-sistent judgments for execution-equivalent code. These findings reveal two signal types with orthogonal errors: execution signals(deterministic but testbench coverage limited)and reasoning signals (semantically rich but hallucination-prone). Their orthog-onality suggests combining the two signals, yet in our experiments letting the reasoner directly observe execution results merely anchors its judgments on test outcomes; we therefore acquire the two signals independently and fuse them only at the decision stage. Based on these insights, we propose EAHC, an Execution-Anchored Hallucination Calibration reranking framework that anchors reasoning judgments to execution behavior so that execution-equivalent candidates receive consistent scores, which implements a dual-channel architecture: EAHC-R, a 4B reasoning discriminator; and EAHC-T, a testbench generator leveraging RAG for execution verification.