Type-IV代码克隆检测:基于分层非对比表示学习
Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning
浏览论文内容
中文总结 AI 辅助
针对Type-IV代码克隆检测,提出基于VICReg框架的非对比分层表示学习方法LWVIC4Code,通过跨层一致性正则化与深度加权生成鲁棒表示,在多个数据集上取得优于或媲美对比学习及大模型的性能。
中文摘要 AI 辅助
软件克隆是彼此相似或功能等价的代码片段,它们对维护、重构和缺陷检测构成重大挑战。检测Type-IV克隆(语义等价但语法可能不同)对于传统的基于词法或语法的方法尤为困难。最近的机器学习方法依赖对比学习,这需要仔细的负采样且可能引入偏差。本文提出LWVIC4Code,一种专门为Type-IV克隆检测设计的非对比表示学习方法。基于方差-不变性-协方差正则化(VICReg)框架和先前的分层VICReg训练,LWVIC4Code引入了跨层一致性正则化和基于深度的层权重,以在Transformer各层间逐步细化语义信息,生成鲁棒且具有判别力的代码表示。我们进行了实证研究,将LWVIC4Code与对比学习基线及零样本大型语言模型在Python(Kamino)和多语言(GPTCloneBench)数据集上进行比较。结果表明,LWVIC4Code在无需负样本的情况下达到具有竞争力或更优的性能,受益于分层监督,并能有效地从Python泛化到其他语言,尤其是Java和C#。这些结果证明,非对比、分层的表示学习是鲁棒语义代码克隆检测的一个有前景的方向。
英文摘要
Software clones are fragments of code that are similar or functionally equivalent to each other. They pose significant challenges for maintenance, refactoring, and bug detection. Detecting Type-IV clones, which are semantically equivalent but may differ syntactically, is particularly difficult for traditional token- or syntax-based methods. Recent machine learning approaches rely on contrastive learning, which requires careful negative sampling and can introduce bias. In this paper, we propose LWVIC4Code, a non-contrastive representation learning approach specifically designed for Type-IV clone detection. Building on the Variance-Invariance-Covariance Regularization (VICReg) framework and prior layer-wise VICReg training, LWVIC4Code introduces cross-layer consistency regularization and depth-dependent layer weighting to progressively refine semantic information across transformer layers, producing robust and discriminative code representations. We conduct an empirical study comparing LWVIC4Code against a contrastive learning baseline and zero-shot large language models on Python (Kamino) and multi-language (GPTCloneBench) datasets. Results show that LWVIC4Code achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and C#. These results demonstrate that non-contrastive, layer-wise representation learning is a promising direction for robust semantic code clone detection.
发表机构
- Université de Montréal(蒙特利尔大学)
机构由 AI 辅助整理,请以论文原文为准。