发表机构
Peking University; Institute for Advanced Algorithms Research; OriginHub Technology(北京大学; 先进算法研究院; 源创科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究LLMs作为训练数据准备者的表现,引入DataPrep - Bench统一基准,涵盖数据构建与质量评估,在多领域和模型上评估,发布相关智能体和评估器,为衡量LLM驱动数据准备能力提供框架。
AI 中文摘要
训练数据的质量从根本上决定了大语言模型(LLMs)的能力,但目前尚无统一基准来衡量LLMs、智能体和以数据为中心的工作流程在端到端准备训练数据方面的表现。我们将LLM驱动的数据准备视为包括数据构建和数据质量评估这两个互补能力。我们引入了DataPrep - Bench,这是首个统一基准,在六个领域和多个基础模型上,根据共享的下游协议联合评估这两种能力。对于数据构建,方法消耗相同原始源并通过微调进行评分;还发布了Data - Construction - Skill智能体。对于数据质量评估,评分函数通过与下游性能的皮尔逊相关性评分,我们发布了基于分布的评估器DAS。DataPrep - Bench为衡量这两种能力的进展提供了统一框架。
英文摘要
The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM-driven data preparation as comprising two complementary capabilities: data construction, which transforms raw sources into supervised training data, and data quality evaluation, which predicts the training value of candidate datasets before downstream training; throughout, "quality" refers to downstream training utility rather than surface-level textual properties. We introduce DataPrep-Bench, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains and multiple base models. For data construction, methods consume identical raw sources and are scored by fine-tuning a base model on their outputs jointly with Dolly-15k; alongside this track we release Data-Construction-Skill, a skill-guided agent that lifts the Dolly-only baseline by nearly 20 points absolute on Llama-3.1-8B Finance and is competitive with the strongest agent- and DataFlow-based methods in knowledge-extraction-dense domains. For data quality evaluation, scoring functions are scored by Pearson correlation with downstream performance on a shared candidate pool; we release the Distributional Alignment Score (DAS), a distribution-based evaluator that uses MMD between a candidate dataset and a domain proxy. DAS attains the strongest cross-model correlation in four of six domains and is the only metric clearing r > 0.70 simultaneously in Math, Science, and Medical, outperforming existing quality-, diversity-, and heuristic-based evaluators. DataPrep-Bench provides a unified, downstream-grounded framework for measuring progress on both capabilities as co-equal targets of LLM-driven data preparation.