arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29283cs.ARcs.LG

RTLCurator:用于RTL生成的标签高效数据整理

RTLCurator: Label-Efficient Data Curation for RTL Generation

Siyang Cai, Cangyuan Li, Wenjing Chang, Kun Wang, Haoyu Gao, Yinhe Han, Ying Wang

首次发表
浏览论文内容

中文总结 AI 辅助

RTLCurator通过对比规范与模拟失败的实现学习兼容性先验,用少量验证配对校准后,平衡对齐度、覆盖度与结构丰富度保留80%语料,在CodeV和RTLCoder上提升RTL生成模型性能,仅需验证10%语料池。

中文摘要 AI 辅助

训练大型语言模型(LLM)编写寄存器传输级(RTL)代码需要大量配对的规范与代码语料,但这类数据十分稀缺,目前多数公开语料都是合成的。合成技术虽能提供规模,但无法保证正确性;在两个广泛使用的RTL数据集中,仅有24.4%和53.5%的配对能通过生成的功能测试。这引发了一个问题:应保留语料中的多少内容,以及保留哪部分?仅靠正确性是不够的:一个在某个边界情况中出错的配对仍能展示有效的语法和接口约定,且复杂的时序设计既更难生成也更难验证,因此按正确性过滤会留下由简短简单模块组成的语料。正确性也难以获取,因为RTL中的行为在表面几乎没有痕迹,验证整个语料仅能将配对分为通过和未通过两类。我们提出RTLCurator,它通过将每个规范与模拟失败的实现进行对比,学习到一种感知行为的兼容性先验,并利用少量已验证的配对将其校准到新的语料中;随后通过平衡对齐度、表示覆盖范围和RTL结构丰富度来构建保留子集。在CodeV和RTLCoder上,以这种方式保留80%的语料,在所有报告的指标上均优于使用完整语料训练的结果,且仅需验证10%的语料池;而仅按分数排序的效果低于随机选择,通过模拟过滤整个语料池的效果也并无提升。

英文摘要

Training large language models (LLMs) to write register-transfer level (RTL) requires large corpora of paired specifications and code, and such data is scarce enough that most public corpora are now synthesized. Synthesis provides scale but not correctness, and in two widely used RTL datasets only 24.4% and 53.5% of pairs pass generated functional tests. This raises the question of how much of such a corpus to keep and which part of it. Correctness alone is a poor answer. A pair that misbehaves in one corner case still shows valid syntax and interface conventions, and complex sequential designs are both harder to generate and harder to validate, so filtering by correctness leaves a corpus of short and simple modules. Correctness is also hard to obtain, since behavior leaves little trace on the surface in RTL, and validating an entire corpus only sorts pairs into passed and failed. We present RTLCurator, which learns a behavior-aware compatibility prior by contrasting each specification with implementations that fail simulation, and calibrates it to a new corpus using a small number of validated pairs. It then constructs the retained subset by balancing alignment, representation coverage, and RTL structural richness. On CodeV and RTLCoder, keeping 80% of the corpus this way improves on training with the full corpus across all reported metrics while validating only 10% of the pool, whereas ranking by the score alone falls below random selection and filtering the whole pool by simulation does no better.

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)
  • Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
  • Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)

机构由 AI 辅助整理,请以论文原文为准。

↑