arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22383cs.SE

通过潜在图学习学习代码的谱表示以实现可泛化的跨语言代码克隆检测

Learning Spectral Representations of Code through Latent Graph Learning for Generalizable Cross-Language Code Clone Detection

Mohsen Hesamolhokama, Ali Sadeghi, Kousha Moeini, Behnam Rohani, Mohammadamin Fazli, Jafar Habibi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出SPECTRA-Siam模型,通过潜在图学习生成跨语言可比的代码谱表示,在多个基准测试中显著提升了代码克隆检测性能,且泛化能力优于基线方法。

中文摘要 AI 辅助

当前代码克隆检测(CCD)方法依赖于固定的、特定语言的图表示,如抽象语法树(AST)或程序依赖图(PDG)。由于功能相同的代码片段可能产生差异极大的结构,这些刚性图会产生无区分度的谱,其性能接近随机水平。为解决该问题,我们提出SPECTRA-Siam,一种孪生潜在图学习网络,通过优化下游CCD性能学习潜在空间,使图的谱成为代码功能的区分性特征。给定代码片段的AST和数据依赖,SPECTRA-Siam通过软槽分配和多头注意力生成固定大小的加权潜在图,并从其归一化拉普拉斯矩阵中提取多尺度谱表示。将所有片段映射到该共享空间可生成跨编程语言的可比谱。在BigCloneBench、AtCoder和四语言CodeNet基准(Java、Python、C++、C#)上的实验支持该设计选择。使用相同的下游分类器,从固定图谱切换到学习的潜在图谱,BigCloneBench上的F1值从0.37提升至0.67,AtCoder上的准确率从0.60提升至0.71。在CodeNet上,完整模型在4个epoch达到0.69准确率,30个epoch后达到0.79。在跨60个未见路径的桥辅助语言迁移中,SPECTRA-Siam的性能仅下降0.058,而基线方法下降0.112至0.228,表明学习的图谱为跨语言克隆检测提供了高度可泛化的表示。

英文摘要

Current code clone detection (CCD) methods rely on fixed, language-specific graph representations like abstract syntax trees (ASTs) or program dependency graphs (PDGs). Because functionally identical code fragments can yield wildly different structures, these rigid graphs produce non-discriminative spectra that perform close to chance. To address this, we propose SPECTRA-Siam, a Siamese latent graph learning network that learns a latent space such that the graph's spectrum serves as a discriminative signature of code functionality by optimizing downstream CCD performance. Given a fragment's AST and data-dependencies, SPECTRA-Siam induces a fixed-size weighted latent graph through soft slot assignment and multi-head attention, and extracts a multi-scale spectral representation from its normalized Laplacian. Mapping all fragments into this shared space yields comparable spectra across programming languages. Experiments on BigCloneBench, AtCoder, and a four-language CodeNet benchmark (Java, Python, C++, C#) support this design choice. Using the same downstream classifier, moving from fixed to learned latent graphs spectra jumps F1 from 0.37 to 0.67 on BigCloneBench and accuracy from 0.60 to 0.71 on AtCoder. On CodeNet, the full model reaches 0.69 accuracy in four epochs and 0.79 after thirty epochs. In bridge-assisted language transfer across 60 unseen paths, SPECTRA-Siam's performance degrades by only 0.058, versus 0.112--0.228 for baselines, showing that learned graph spectra provide a highly generalizable representation for cross-language clone detection.

↑