发表机构
Google(谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型数据质量问题,引入可扩展的基于影响的数据审核管道,通过映射语义邻域到有向图,利用参考LLM概率分布评估数据效用,转换为局部优势指标,有效清理数据集,揭示基准漏洞,确保数据完整性。
AI 中文摘要
大语言模型(LLM)的对齐越来越受到数据质量的瓶颈限制。随着数据集规模的扩大,大量偏好和指令微调语料库不可避免地积累了隐藏的结构矛盾、安全风险和系统性人工标注错误。标准的数据集审核方法,如语义重复数据删除或使用LLM作为评判工具,难以捕捉单个记录的实际预测影响,并且常常错过深层功能规则冲突。为了解决这个问题,我们引入了一种可扩展的、仅推理的数据评估管道,该管道无需迭代模型重新训练即可近似夏普利值。通过将语义k近邻邻域映射到有向图中,我们的框架使用零样本和单样本条件对数似然转移,直接通过参考LLM的概率分布评估数据效用。然后,我们的管道将这些预测影响分数转换为局部优势指标,以隔离梯度冲突记录。我们展示了该管道在清理两个经过严格审查的对齐数据集方面的有效性。首先,将我们的管道应用于HelpSteer2数据集,将人工审核搜索空间减少了99.1%,成功发现了各种故障模式下的错误标注记录。其次,将我们的自动审核策略应用于Anthropic的HH-RLHF训练和评估分割,识别出数千个隐藏的安全和事实偏好反转。至关重要的是,通过将此审核扩展到评估分割,我们揭示了当前基准完整性中的严重漏洞:高能力模型经常预测更安全或更有帮助的响应,但却受到客观上有缺陷的人工地面真值标签的惩罚。总体而言,我们的工作提供了一种基于数学的、高效的诊断工具,以发现人工标签故障、清理评估基准并确保LLM对齐数据的完整性。
英文摘要
The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks, and systemic human annotation errors. Standard dataset auditing methods, such as semantic deduplication or LLM-as-a-judge, struggle to capture the actual predictive impact of individual records and often miss deep functional rule clashes. To address this, we introduce a scalable, inference-only data valuation pipeline that approximates the Shapley value without iterative model retraining. By mapping semantic k-NN neighborhoods into a directed graph, our framework evaluates data utility directly through a reference LLM's probability distribution using zero-shot and one-shot conditional log-likelihood shifts. Our pipeline then translates these predictive influence scores into localized advantage metrics to isolate gradient-conflicting records. We demonstrate the pipeline's efficacy in sanitizing two heavily vetted alignment datasets. First, applying our pipeline to the HelpSteer2 dataset reduced the manual audit search space by 99.1%, successfully uncovering falsely-labeled records across diverse failure modes. Second, applying our automated audit strategy to Anthropic's HH-RLHF training and evaluation splits identified thousands of hidden safety and factual preference inversions. Crucially, by extending this audit to the evaluation split, we expose severe vulnerabilities in current benchmark integrity: highly capable models frequently predict the safer or more helpful response, only to be penalized by objectively flawed human ground-truth labels. Overall, our work provides a mathematically grounded, highly efficient diagnostic tool to uncover human label failures, sanitize evaluation benchmarks, and ensure the integrity of LLM alignment data.