arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32987stat.MEmath.STstat.MLstat.TH

迭代合成数据增强下实体排序的最大似然估计

Maximum Likelihood Estimation for Entity Ranking under Iterative Synthetic Data Augmentation

Yiqiao Jin, Mengxin Yu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对BTL模型下稀疏比较数据的实体排序问题,分析了迭代合成数据增强对最大似然估计的影响,推导了最优统计速率与渐近正态性,并刻画了避免模型崩溃的条件,通过实验验证了理论结果。

中文摘要 AI 辅助

我们研究了在Bradley--Terry--Luce (BTL)模型下,利用稀疏比较图收集的成对比较数据进行实体排序和推断的问题。在实践中,收集高质量的人类判断可能既昂贵又耗时,这促使在模型训练中使用合成数据增强。我们分析了一种迭代合成增强工作流,在该工作流中,从拟合模型生成的合成比较被连续添加到原始数据集中。在此过程中,随着迭代次数的增加,真实数据的比例可能会消失。然而,新兴文献表明,对合成数据进行递归训练可能导致模型崩溃,这引发了对这类增强过程统计可靠性的担忧。为此,我们在有限样本、高维场景下系统分析了由此产生的最大似然估计。对于得到的迭代最大似然估计量(MLE),我们推导了其最优的有限样本$\u2113_2$和$\u2113_{\infty}$统计速率,并在自然的可识别性条件下建立了其渐近正态性。我们进一步刻画了尽管真实数据比例不断减少但仍能避免模型崩溃的场景。我们通过大规模数值实验和对Arena Human Preference 140k数据集的应用验证了我们的理论发现。

英文摘要

We study entity ranking and inference under the Bradley--Terry--Luce (BTL) model using pairwise comparisons collected over sparse comparison graphs. In practice, collecting high-quality human judgments can be expensive and time-consuming, motivating the use of synthetic data augmentation in model training. We analyze an iterative synthetic augmentation workflow in which synthetic comparisons generated from fitted models are successively added to the original dataset. In this process, the proportion of real data may vanish as the number of iterations grows. However, emerging literature has shown that recursive training on synthetic data can lead to model collapse, raising concerns about the statistical reliability of such augmentation procedures. To this end, we systematically analyze the resulting MLE in finite-sample, high-dimensional regimes. For the resulting iterative maximum likelihood estimator (MLE), we derive its optimal finite-sample $\ell_2$ and $\ell_{\infty}$ statistical rates and establish its asymptotic normality under natural identifiability conditions. We further characterize regimes in which model collapse is avoided despite the diminishing fraction of real data. We validate our theoretical findings through large-scale numerical experiments and an application to the Arena Human Preference 140k dataset.

发表机构

  • Washington University in St. Louis(华盛顿大学圣路易斯分校)

机构由 AI 辅助整理,请以论文原文为准。

↑