发表机构
China University of Mining and Technology; Macau University of Science and Technology(中国矿业大学; 澳门科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RAPO-Sol两阶段训练框架,结合检索增强微调和基于语义扰动的直接偏好优化,在仓库级Solidity代码生成中显著提升性能,且推理时无需检索。
AI 中文摘要
用Solidity编写的智能合约管理资产、权限和不可逆的状态变更,这使得代码生成既实用又对安全性至关重要。仓库级Solidity生成具有挑战性,因为模型必须综合完整的合约或库,同时保持状态变量、修饰符、事件、继承、外部调用和访问控制逻辑之间的一致性。我们提出了RAPO-Sol,一个用于仓库级Solidity代码生成的两阶段训练框架。首先,检索增强微调(RAFT)用相似的Solidity示例增强每个训练输入,帮助模型学习重复出现的合约级模式,同时在推理时保持无需检索。其次,直接偏好优化(DPO)训练模型偏好参考合约,而非接近但在语义上有缺陷的替代方案。我们使用Solidity语义锚点扰动(SAP)构建被拒绝的样本,该扰动会扰动验证语句、可见性修饰符、数据位置关键字、上下文变量、支付操作和低级调用。在SolidityBench上使用CodeLlama-7B-Instruct、DeepSeek-Coder-6.7B-Instruct和Qwen2.5-Coder-7B-Instruct进行的实验表明,RAFT持续优于监督微调,而基于SAP的DPO在BLEU和SolidityScore上提供了进一步的提升。完整的RAFT+DPO流程在所有三个模型上均取得了最佳性能,展示了训练期间检索和Solidity感知偏好优化的互补优势,而无需在推理时增加检索成本。
英文摘要
Smart contracts written in Solidity manage assets, permissions, and irreversible state changes, making code generation both useful and security-critical. Repository-level Solidity generation is challenging because models must synthesize complete contracts or libraries while preserving consistency across state variables, modifiers, events, inheritance, external calls, and access-control logic. We present RAPO-Sol, a two-stage training framework for repository-level Solidity code generation. First, Retrieval-Augmented Fine-Tuning (RAFT) augments each training input with similar Solidity examples, helping the model learn recurring contract-level patterns while remaining retrieval-free at inference time. Second, Direct Preference Optimization (DPO) trains the model to prefer reference contracts over close but semantically flawed alternatives. We construct rejected samples using Solidity Semantic-Anchor Perturbation (SAP), which perturbs validation statements, visibility modifiers, data-location keywords, context variables, payment operations, and low-level calls. Experiments on SolidityBench with CodeLlama-7B-Instruct, DeepSeek-Coder-6.7B-Instruct, and Qwen2.5-Coder-7B-Instruct show that RAFT consistently improves over supervised fine-tuning, while SAP-based DPO provides further gains in BLEU and SolidityScore. The full RAFT+DPO pipeline achieves the best performance across all three models, demonstrating complementary benefits from retrieval during training and Solidity-aware preference optimization without adding retrieval cost at inference.
Comments12 pages, 3 figures, 5 tables