发表机构
Purdue University; Meta(普渡大学; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出FROST在线框架,利用真实数据锚定的梯度反馈估计合成数据效用,动态过滤噪声样本,在图像分类、文本到SQL及工业广告重排序中提升性能。
AI 中文摘要
当真实世界数据有限时,合成数据可以扩展训练监督,但噪声和分布不匹配会降低其价值。现有的合成数据选择方法往往强调保真度或多样性,而非学习者不断变化的需求。我们提出FROST,一种在线框架,通过锚定在真实训练数据中的梯度反馈来估计合成数据的效用。它根据近期历史校准批次效用以确定何时需要过滤,并且仅在带外批次中过滤样本以确定保留什么,无需外部验证器或保留验证集。在两个公共基准上的图像分类和文本到SQL的LLM微调实验中,FROST过滤掉约20-30%的合成数据,同时相比在完整合成数据池上训练提高了真实任务性能。我们进一步在大规模工业广告重排序系统的训练中应用FROST,在高度优化的生产基线上取得了显著的性能提升,证明了其有效性和泛化能力。
英文摘要
Synthetic data can scale training supervision when real-world data are limited, but noise and distribution mismatch can reduce its value. Existing synthetic data selection methods often emphasize fidelity or diversity rather than the learner's evolving needs. We propose FROST, an online framework that estimates synthetic-data utility through gradient feedback anchored in real training data. It calibrates batch utility against recent history to determine when filtering is needed and filters samples only in out-of-band batches to determine what to retain, without an external verifier or held-out validation set. Experiments on two public benchmarks for image classification and LLM fine-tuning for text-to-SQL show that FROST filters out around 20--30% of the synthetic data while improving real-task performance compared with training on the full synthetic data pool. We further apply FROST during training in a large-scale industrial ads re-ranking system, achieving significant performance gains over a highly optimized production baseline, demonstrating its effectiveness and generalizability.
Comments21 pages, 6 figures