arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型的频率感知持续学习智能合约漏洞检测

Frequency-Aware Continual Learning for Smart Contract Vulnerability Detection with Large Language Models

Tenghui Huang, Jiawen Kang, Dongning Liu, Changyan Yi, Chengjun Cai, Anjia Yang, Li Li, Dong In Kim

arXiv 2608.19680首次发表:更新:

发表机构

School of Automation, Guangdong University of Technology; School of Computer Science and Technology, Guangdong University of Technology; College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics; Department of Computer Science, City University of Hong Kong (Dongguan); College of Cyber Security, Jinan University; Guangdong Institute of Science and Technology Information; Department of Electrical and Computer Engineering, Sungkyunkwan University(广东工业大学自动化学院; 广东工业大学计算机科学与技术学院; 南京航空航天大学计算机科学与技术学院; 香港城市大学(东莞)计算机科学系; 暨南大学网络安全学院; 广东省科技情报研究所; 成均馆大学电气与计算机工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型用于智能合约漏洞检测时的三个挑战,提出含FA-LoRA、FAR、APPM的三阶段集成框架,在DIVE数据集上取得优异性能,可有效应对区块链生态系统的相关问题。

AI 中文摘要

用大语言模型(LLMs)进行智能合约漏洞检测面临三个因果关联的挑战:其一,新漏洞类别需要参数高效适配,因为对依次到达的任务进行全量重训练是不可行的;其二,在共享主干模型上训练每个任务的适配器会导致先前学习到的漏洞发生灾难性遗忘;其三,由于推理时任务身份未知,必须将众多适配器整合为单个模型。每个挑战都直接源于对前一个挑战的解决方案,因此需要一个集成框架。本文提出了一个三阶段流程,每个阶段解决一个挑战并衔接下一阶段:适配阶段采用频率感知低秩适配(FA-LoRA),该方法在傅里叶域中进行适配,使用逐频率重要性门,仅需0.4%的可训练参数,性能优于标准LoRA和QLoRA;持续学习阶段应用遗忘感知重放(FAR),它利用这些频率门通过损失动态估计每个样本的遗忘风险,并优先选择易受攻击的知识进行重放,在连续任务上实现了0.8022的平均微F1值;部署阶段采用锚点保护渐进合并(APPM),它利用FAR训练产生的非对称泛化效果,将泛化能力最强的适配器识别为锚点,并通过带频域门竞争的锚点保护加权合并将所有适配器整合为单个模型。APPM实现了0.8085的微F1值,与独立任务上限的差距在2.7%以内,合并耗时156毫秒,无额外运行时内存开销。在DIVE数据集上的实验证实,该框架能有效应对不断发展的区块链生态系统的这三个挑战。

英文摘要

Smart contract vulnerability detection with Large Language Models (LLMs) faces three causally linked challenges. First, new vulnerability categories demand parameter-efficient adaptation, since full retraining is prohibitive for sequentially arriving tasks. Second, training per-task adapters on a shared backbone causes catastrophic forgetting of previously learned vulnerabilities. Third, the resulting multiplicity of adapters must be consolidated into a single model, since task identity is unknown at inference time. Each challenge arises directly from the solution to its predecessor, making an integrated framework essential. We propose a three-stage pipeline in which each stage addresses one challenge and feeds into the next. The adaptation stage uses Frequency-Aware Low-Rank Adaptation (FA-LoRA), which performs adaptation in the Fourier domain with per-frequency importance gates, requiring only 0.4% trainable parameters while outperforming standard LoRA and QLoRA. The continual learning stage applies Forget-Aware Replay (FAR), which uses these frequency gates to estimate per-sample forgetting risk via loss dynamics and prioritizes vulnerable knowledge for rehearsal, achieving an average Micro-F1 of 0.8022 across sequential tasks. The deployment stage employs Anchor-Protected Progressive Merging (APPM), which exploits the asymmetric generalization produced by FAR training to identify the strongest-generalizing adapter as an anchor and consolidates all adapters into a single model via anchor-protected weighted merging with frequency-domain gate competition. APPM achieves a Micro-F1 of 0.8085, within 2.7% of the independent per-task upper bound, at a merge cost of 156 ms and no additional runtime memory. Experiments on DIVE confirm the framework effectively addresses all three challenges for evolving blockchain ecosystems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑