arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过训练动态刻画合成数据

Synthetic Data Characterization via Training Dynamics

Irene Lago, Ana Ezquerro, David Vilares

arXiv 2609.39447首次发表:更新:

发表机构

Universidade da Coruña; Graz University of Technology(拉科鲁尼亚大学; 格拉茨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过样本级可学习性刻画LLM生成数据,对比不同模型家族和规模,并评估基于该信号的数据选择策略对合成与人类数据的不同影响。

AI 中文摘要

解释LLM生成数据的属性对于理解其在学习任务中的效用和局限性非常重要。在本工作中,我们通过样本级可学习性来刻画合成数据,研究不同LLM家族和规模之间的差异,并以人类撰写的数据作为参考。我们首先生成涵盖单标签和多标签分类、标注及树预测任务的合成数据集。然后,我们从编码器训练动态中推导出机器数据和有机数据的经验数据分布,并估计这些分布在编码器间的稳健性。最后,我们评估基于这些可学习性信号的数据选择策略如何对两种数据源产生不同影响。

英文摘要

Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.

CommentsAccepted at Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑