arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11746cs.LGcs.CL

面向分布外泛化的、受Epiplexity引导的数据选择与生成

Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization

  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes, Andrew Gordon Wilson

AI总结:

该研究提出以Epiplexity为引导,通过数据选择与合成数据生成提升模型分布外泛化能力,实验显示更高Epiplexity可改善零样本及微调任务的下游性能。

AI中文摘要:

现代系统日益需要在训练期间未指定的任务之间进行迁移。在这些新的、未预料到的场景中,什么样的数据有助于泛化?一个假设是,具有更多结构信息的数据可能包含可在更广泛的下游场景中复用的共享回路和子程序。Epiplexity是最近提出的一种衡量受计算约束的学习者可从数据中提取的结构信息的指标,它为推理这种关系提供了机制。本文中,我们展示了如何将Epiplexity实现为用于数据选择和合成数据生成的在线训练信号。对于选择,我们对自然数据域的训练损失曲线拟合缩放定律,以预测作为训练token函数的预期Epiplexity增益,并使用该信号在训练期间自适应确定各域的采样权重。对于合成数据生成,我们将生成器的奖励定义为学习者在先前生成数据的缓冲区内的Epiplexity变化,并使用REINFORCE策略梯度引导生成器朝向Epiplexity最大化的分布。在这两种情况下,更高的Epiplexity都能预测零样本和基于微调的任务上的下游性能提升,支持了富含结构信息的数据能产生可跨域迁移的表示这一假设。

英文摘要:

Modern systems are increasingly expected to transfer across tasks not specified during training. What data facilitates generalization in these new, unanticipated settings? One hypothesis is that data with more structural information could contain shared circuits and subprograms that could be recycled in a wider array of downstream settings. Epiplexity, a recently proposed measure of the structural information a compute-bounded learner can extract from data, provides a mechanism to reason about this relationship. In this paper, we show how to operationalize epiplexity as an online training signal for data selection and synthetic data generation. For selection, we fit scaling laws to the training loss curves of natural data domains to predict the expected epiplexity gain as a function of training tokens, and use this signal to adaptively determine the sampling weights over domains during training. For synthetic data generation, we define a generator's reward as the change in learner epiplexity over a buffer of previously generated data and use REINFORCE policy gradients to guide the generator toward an epiplexity-maximizing distribution. In both cases, higher epiplexity predicts improved downstream performance on zero-shot and fine-tuning based tasks, supporting the hypothesis that data rich in structural information yield representations that transfer across domains.

补充信息

↑