用于鲁棒代码表示的不变预训练
Invariant Pretraining for Robust Code Representations
浏览论文内容
中文总结 AI 辅助
本文针对代码表示模型在不变程序下鲁棒性退化问题,提出仅基于代码的不变预训练(InvPT)方法,在克隆检测、代码分类任务上分别提升鲁棒性最高11、19个百分点,同时保持或提升标准准确率。
中文摘要 AI 辅助
基于编码器的代码表示模型被广泛应用于克隆检测、代码分类等判别任务,其体积小、推理成本低是关键优势。然而这类模型的鲁棒性十分脆弱:在不变程序(即语义等价但语法形式不同的代码)下,即便程序行为未改变,学习到的表示也会大幅退化。本文针对四个编码器基线模型、两个下游任务及四个数据集开展了关于该鲁棒性差距的实证研究,并提出了一种仅基于代码的极简持续预训练方案,可大幅缩小该差距。本文的方法为不变预训练(InvPT),它对语料库应用语义保留变换,并将掩码语言建模与多正样本监督对比学习相结合,将同一源函数的所有增强版本视为正样本,将自对比对(同一代码、不同掩码)与不变对比对(变换后的代码)混合,以获得不同难度的正样本。与现有的对比式代码编码器不同,InvPT无需成对的自然语言数据。在评估中,InvPT在变换后的测试集上,克隆检测任务的鲁棒性提升最高达11个百分点,代码分类任务提升最高达19个百分点,同时匹配或提升了标准准确率; ablation实验表明多正样本不变对比是性能提升的主要来源。本文的目标并非提出新的目标函数,而是仔细探究编码器鲁棒性失效的场景,以及一种简单的仅基于代码的方案能在多大程度上恢复该鲁棒性。
英文摘要
Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive. Their robustness, however, is fragile: under invariant programs, semantically equivalent code written in different syntactic forms, learned representations degrade substantially even though program behavior is unchanged. We present an empirical study of this robustness gap across four encoder baselines, two downstream tasks, and four datasets, together with a minimal code-only continued pretraining recipe that closes much of it. Our method, invariant pretraining (InvPT), applies semantic-preserving transformations to the corpus and combines masked language modeling with multi-positive supervised contrastive learning that treats all augmentations of the same source function as positives, mixing self-contrast pairs (same code, different masks) with invariant-contrast pairs (transformed code) for positives of varying difficulty. Unlike prior contrastive code encoders, InvPT does not require paired natural-language data. Across our evaluation, InvPT improves robustness on transformed test sets by up to 11 percentage points on clone detection and 19 on code classification while matching or improving standard accuracy, and our ablations isolate multi-positive invariant contrast as the main source of the gains. Our aim is not a new objective but a careful measurement of where encoder robustness breaks and how far a simple, code-only recipe can recover it.
发表机构
- University of California at Davis(加州大学戴维斯分校)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。