发表机构
University of Saskatchewan(萨斯喀彻温大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出通过模型合并技术构建跨域代码克隆检测器,经实验验证TIES合并方法泛化性更优,合并后检测器性能优于零样本代码大模型且泛化性更强。
AI 中文摘要
代码克隆类型日益多样化,从语法拷贝到跨语言语义克隆再到AI生成的重复代码,已引发克隆检测的碎片化危机。当前深度学习检测器均为领域专家模型,在训练分布外性能大幅下降,跨领域F1值降幅超过70%。部署多个专用模型不切实际,而训练单个跨域检测器需要同时访问所有训练数据。为解决该问题,本文研究仅基于训练后检查点运行的事后技术——模型合并,评估了五种任务向量方法的参数合并、贪心层拼接的架构合并,以及四种代码模型、三个基准、十二种配置下的跨分词器对齐。同基TIES合并可生成有效的跨域检测器,在两个模型族和三个随机种子上得到验证,UniXcoder上的组合F1值达0.865,合并步骤无需任何训练数据,即可达到多任务性能的93%。WUDI在分布内组合F1值最高,为0.899,但TIES对未见过的AI生成克隆泛化性更好,成为本文推荐方法。跨基合并在五种方法中仅产生微小且方差大的增益,表明通过共享预训练基的任务向量兼容性是有效合并的关键因素。合并后的检测器在GPTCloneBench上的推理成本更低,且对未见过的AI生成克隆的泛化性比多任务训练高4倍,优于零样本代码大模型,说明存在域内性能与分布外鲁棒性的权衡。本文是软件工程领域首批针对模型合并的系统实证研究之一,为构建跨域克隆检测器提供了实用方案。
英文摘要
The growing diversity of code clone types, from syntactic copies to cross-language semantic clones to AI-generated duplicates, has created a fragmentation crisis in clone detection. Current deep learning detectors are domain specialists that degrade significantly outside their training distribution, with F1 drops exceeding 70% across domains. Deploying multiple specialized models is impractical, yet training a single cross-domain detector requires simultaneous access to all training data. To address this, we investigate model merging, a family of post-hoc techniques that operate solely on trained checkpoints. We evaluate parameter merging with five task-vector methods, architecture merging via greedy layer stitching, and cross-tokenizer alignment across four code models, three benchmarks, and twelve configurations. Same-base TIES merging creates effective cross-domain detectors, validated across two model families and three random seeds, reaching 0.865 combined F1 on UniXcoder, 93% of multi-task performance without any training data at the merging step. WUDI achieves the highest in-distribution combined F1 at 0.899, but TIES generalizes better to unseen AI-generated clones, making it our recommended method. Cross-base merging yields only marginal and high-variance gains across all five methods, indicating that task vector compatibility through a shared pre-trained base is the binding factor for effective merging. Merged detectors also outperform zero-shot code LLMs on GPTCloneBench at lower inference cost and generalize up to 4x better than multi-task training to unseen AI-generated clones, suggesting a trade-off between in-domain performance and OOD robustness. This work provides one of the first systematic empirical studies of model merging for software engineering and a practical recipe for building cross-domain clone detectors.
CommentsAccepted at ASE 2026
Journal refProceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE '26), October 12--16, 2026, Munich, Germany