保持评估公平:通过成员推断攻击检测代码生成基准中的数据泄漏
Keep Evaluation Fair: Detecting Data Leakage in Code Generation Benchmarks via Membership Inference Attacks
浏览论文内容
中文总结 AI 辅助
针对代码生成基准数据泄漏问题,提出CGMIA成员推断攻击方法,结合多种特征检测泄漏,在八个基准上优于现有方法。
中文摘要 AI 辅助
代码生成基准被广泛用于评估大型语言模型(LLMs),但基准数据泄漏到训练集中会虚增性能并削弱评估的有效性。DetectLeak是一种专门为代码生成基准泄漏检测设计的方法,它依赖困惑度分数来识别可能泄漏的样本。然而,困惑度主要反映对代码模式的普遍熟悉程度,在复杂或罕见样本上可能表现不佳。它还忽略了其他有用的信号,如代码相似性、功能正确性和语义表示。为解决这些局限性,我们提出了CGMIA(代码生成特定成员推断攻击),一种用于检测代码生成基准中泄漏的方法。CGMIA在基准样本子集上微调一个影子模型,以构建有标签的成员和非成员数据。对于每个样本,它收集输入提示、生成的代码和参考解决方案,并提取专家特征,包括CodeBLEU、编辑距离、测试通过率和困惑度,以及来自CodeBERT嵌入的语义特征。一个集成学习模块结合这些特征,以捕捉表面层面的记忆信号和更深层次的行为模式,使分类器能够预测样本是否包含在目标模型的训练集中。在八个代码生成基准上的实验表明,CGMIA在大多数情况下优于八种现有的成员推断方法。它还能有效检测StarCoder-7B训练数据中已知泄漏的APPS样本。
英文摘要
Code generation benchmarks are widely used to evaluate Large Language Models (LLMs), but benchmark data leakage into training sets can inflate performance and undermine evaluation validity. DetectLeak, a method specifically designed for code generation benchmark leakage detection, relies on perplexity scores to identify likely leaked samples. However, perplexity mainly reflects general familiarity with code patterns and may perform poorly on complex or rare samples. It also overlooks other useful signals, such as code similarity, functional correctness, and semantic representations. To address these limitations, we propose CGMIA (Code-Generation-specific Membership Inference Attack), a method for detecting leakage in code generation benchmarks. CGMIA fine-tunes a shadow model on a subset of benchmark samples to construct labeled member and non-member data. For each sample, it collects the input prompt, generated code, and reference solution, and extracts expert features, including CodeBLEU, edit distance, test pass rate, and perplexity, together with semantic features from CodeBERT embeddings. An integrated learning module combines these features to capture both surface-level memorization signals and deeper behavioral patterns, enabling a classifier to predict whether a sample was included in the target model's training set. Experiments on eight code generation benchmarks show that CGMIA outperforms eight existing membership inference methods in most cases. It also effectively detects known leaked APPS samples in StarCoder-7B's training data.
发表机构
- School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院)
- Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)
- Engineering Research Center of Transportation Information and Safety (ERCTIS), MoE of China, Wuhan University of Technology(武汉理工大学交通运输部交通信息与安全工程研究中心)
- State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。