arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22765cs.AR

质量优于数量:面向高效Verilog代码生成的多样性感知数据选择

Quality over Quantity: Diversity-Aware Data Selection for Efficient Verilog Code Generation

  • Northwestern Polytechnical University(西北工业大学)
  • Nantong Normal College(南通师范高等专科学校)
  • East China Normal University(华东师范大学)
  • Nantong University(南通大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Yiheng Shen, Wei Zheng, Xiao Wei, Hao Shen, Xiang Chen, Guang Yang

AI总结:

针对Verilog代码生成中数据质量与多样性不足的问题,提出VeriSelector框架,联合优化质量与多样性,仅用20%-25%数据即超越全量训练,显著提升Pass@1并减少80%训练时间。

AI中文摘要:

大型语言模型(LLMs)在Verilog代码生成方面展现出巨大潜力,然而现有数据集包含大量噪声和冗余。先前的数据选择方法仅处理孤立的质量方面,忽视了训练集的全局多样性,并且无法捕捉Verilog特有的结构语义。为弥补这一差距,我们提出了VeriSelector,这是首个针对Verilog代码生成的数据选择框架,它联合优化质量和多样性。我们将选择问题形式化为一个受约束的双目标子集选择问题,并通过三阶段近似求解。在质量方面,一个多粒度流水线首先通过测试台仿真验证功能正确性,然后通过指令遵循难度(IFD)评分过滤不匹配的样本。在多样性方面,109维的Verilog特有结构特征(AST、CFG和网表)与文本嵌入融合,用于基于聚类的多样性建模。随后,一种比例自适应采样策略根据IFD排名分配每簇配额,并具有可证明的分布保持保证。在三个LLM和三个基准上的实验表明,VeriSelector仅使用20%-25%的数据即可超越全数据集训练和最先进的基线,实现超过118%的性能保持率,并将训练时间减少80%以上。值得注意的是,在所有评估模型中,VeriSelector相比全数据集训练将平均Pass@1提高了18.49%-29.43%,相比表现最佳的基线提高了1.36%-7.95%。

英文摘要:

Large Language Models (LLMs) have shown remarkable potential in Verilog code generation, yet existing datasets contain con siderable noise and redundancy. Prior data selection methods address only isolated quality aspects, neglect the global diversity of the training set, and cannot capture Verilog-specific structural semantics. To bridge this gap, we propose VeriSelector, the first data selection framework for Verilog code generation that jointly optimizes quality and diversity. We formulate the selec tion problem as a constrained bi-objective subset selection problem and solve it via a three-stage approximation. For quality, a multi-granularity pipeline first verifies functional correctness through testbench simulation and then filters misaligned samples via Instruction-Following Difficulty (IFD) scoring. For diversity, 109-dimensional Verilog-specific structural features (AST, CFG, and Netlist) are fused with textual embeddings for clustering-based diversity modeling. A proportional adaptive sampling strat egy then allocates per-cluster quotas guided by IFD ranks, with a provable distribution preservation guarantee. Experiments on three LLMs and three benchmarks show that VeriSelector outperforms full-dataset training and state-of-the-art baselines us ing only 20%-25% of the data, achieving Performance Retention Rates above 118% and reducing training time by over 80%. Notably, VeriSelector improves average Pass@1 by 18.49%-29.43% over full-dataset training and by 1.36%-7.95% over the best-performing baseline across all evaluated models.

补充信息

↑