校准用于大语言模型数据污染检测的训练后特征偏移
Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection
浏览论文内容
中文总结 AI 辅助
针对训练后处理导致的LLM数据污染检测中特征偏移问题,提出CalibDCD校准框架,经实验可提升现有检测器的AUC和TPR@5%FPR指标。
中文摘要 AI 辅助
大语言模型(LLMs)在海量且大多未公开的语料库上训练,这些语料库可能包含受版权保护或涉及隐私的内容。因此,数据污染检测(DCD)旨在确定给定文本是否为目标LLM预训练语料库的成员。近期的最优DCD方法遵循基于特征的范式,该范式从输入文本和对应的模型输出中提取成员特征。然而,大多数现代LLM会经历训练后处理,如指令调优、偏好优化和面向推理的训练,这会改变模型输出并偏移对应的成员特征,从而降低成员与非成员之间的可分性。为解决此问题,我们提出CalibDCD,这是一个适用于基于特征的DCD方法的通用校准框架,包含:(1)多视图偏移检测,用于识别与训练后处理相关的重复特征偏移;(2)有界特征校正,用于选择性减轻这些偏移对成员预测的影响。具体而言,多视图偏移检测在已知非成员文本上评估受控提示变体,并整合最具信息性的视图以识别重复特征偏移;有界特征校正选择性调整与检测到的偏移对齐的特征分量,并控制校正范围以保留有用的检测信息。实验表明,CalibDCD可持续改进现有基于特征的检测器,在AUC上的提升最高达7.0%,在TPR@5%FPR上的提升最高达15.0%。
英文摘要
Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) therefore aims to determine whether a given text is a member of the pre-training corpus of a target LLM. Recent state-of-the-art DCD methods follow a feature-based paradigm that derives membership features from the input text and the corresponding model output. However, most modern LLMs undergo post-training, such as instruction tuning, preference optimization, and reasoning-oriented training, which can alter model outputs and shift the corresponding membership features, thereby reducing the separability between members and non-members. To address this problem, we propose CalibDCD, a broadly applicable calibration framework for feature-based DCD methods, comprising (1) Multi-View Shift Detection, which identifies recurring feature shifts associated with post-training, and (2) Bounded Feature Correction, which selectively mitigates their influence on membership prediction. Specifically, Multi-View Shift Detection evaluates controlled prompt variants on known non-member texts and consolidates the most informative views to identify recurring feature shifts. Bounded Feature Correction selectively adjusts feature components aligned with the detected shifts and controls the correction extent to preserve useful detection information. Experiments show that CalibDCD consistently improves existing feature-based detectors, with gains of up to 7.0% in AUC and 15.0% in TPR@5%FPR.
发表机构
- The University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。