SynthSentry:检测语言模型训练数据中的合成数据污染
SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data
浏览论文内容
中文总结 AI 辅助
SynthSentry提出一种模型无关的语料级污染信号,通过词汇多样性、n-gram尾部及困惑度方差检测合成数据,实现训练前筛选,并验证了校准与剪枝风险。
中文摘要 AI 辅助
在自身或其他模型输出上递归训练的大型语言模型会经历模型崩溃,此时分布尾部与事实准确性恶化,而流畅性得以保留。先前的工作在训练后诊断崩溃;可操作的问题是在训练前筛选来源未知的语料库。我们引入了SynthSentry,一种语料库级别、模型无关的污染信号,无需访问生成模型、无需生成历史、也无需合成标签。该得分是基于三个统计量的分布散度:词汇多样性崩溃、n-gram尾部截断以及跨参考模型的困惑度方差。我们在由小型开放权重生成器和指令微调的开放权重模型污染的语料库上,采用留一生成器协议进行评估。一项按领域分层的研究测量了自然重复人类文本(法律、临床、源代码)上的假阳性率。该得分按严重程度对语料库进行排序,当整个生成器家族被排除时,损失很小。一旦使用协方差收缩和自举阈值替代朴素分位数(其运行超出预算四倍),各领域的校准接近其名义假阳性预算。下游微调检查显示,在我们的规模下没有污染驱动的准确性缺陷,因此剪枝是否能恢复准确性仍是一个开放问题;同一运行显示,一旦剪枝超过真实污染比例,存在过度剪枝风险。我们将筛选视为数据整理防御而非事后诊断,并发布了评分工具包。所有结果均为小规模;范围是英语、批处理模式的语料库筛选。污染源是单代或手工编写的,而非递归生成的,因此结果适用于一般合成污染,而非递归深度。
英文摘要
Large language models trained recursively on their own or other models' outputs undergo model collapse, in which distributional tails and factual accuracy deteriorate while fluency survives. Prior work diagnoses collapse after training; the actionable problem is screening a corpus of unknown provenance before training. We introduce SynthSentry, a corpus-level, model-agnostic contamination signal requiring no access to the generating model, no generation history, and no synthetic labels. The score is a distributional divergence over three statistics: lexical diversity collapse, n-gram tail truncation, and perplexity variance across reference models. We evaluate on corpora contaminated by small open-weight generators and an instruction-tuned open-weight model under a leave-one-generator-out protocol. A domain-stratified study measures false positives on naturally repetitive human text (legal, clinical, source code). The score ranks corpora by severity with little loss when whole generator families are held out. Per-domain calibration holds near its nominal false-positive budget once covariance shrinkage and a bootstrap threshold replace a naive quantile, which runs four times over budget. A downstream fine-tuning check showed no contamination-driven accuracy deficit at our scale, so whether pruning recovers one remains open; the same run shows over-pruning risk once pruning exceeds the true contamination fraction. We frame screening as a data-curation defense rather than a post-hoc diagnosis and release the scoring toolkit. All results are small-scale; scope is English-language, batch-mode corpus screening. Contamination sources are single-generation or hand-authored rather than recursively generated, so results speak to synthetic contamination generally and not to recursion depth.