我应该使用这个合成数据集进行训练吗?如何用最少的真实数据进行测试
Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real Data
浏览论文内容
中文总结 AI 辅助
本文提出aeSFT方法,用极少真实数据判断合成数据集是否可用于训练,实验表明其识别有用合成数据的效率优于均值序贯测试,性能与固定样本符号翻转测试等相当且假阳性率达标。
中文摘要 AI 辅助
数字孪生(DTs)和学习到的世界模型越来越多地用于生成合成数据,以补充工程系统中用于训练人工智能(AI)模型的稀缺真实数据集。然而,由于不可避免的仿真到现实(sim-to-real)差距,这种数据增强可能无法提高训练模型在真实数据分布上的性能。本文解决由此产生的决策问题:给定一个真实数据集、一个候选合成数据集和一个固定的学习算法,判断在增强数据集上训练是否能提高真实的总体水平性能,同时消耗尽可能少的真实测试数据点。本文考虑两种公式:一种是对两个训练模型之间的平均损失差异进行直接测试,另一种是对配对损失差异进行基于对称性的测试,后者用更强的零假设换取更快的证据积累。对于后一种测试,本文引入了自适应e过程符号翻转测试(adaptive e-process sign-flip test,aeSFT),这是一种双重自适应程序,可同时调整蒙特卡洛符号翻转轮次的数量(从而调整计算成本)和消耗的真实测试数据量。aeSFT具有随时有效的I类错误控制,无需预先指定测试集大小。在合成数据分类任务、DT辅助的无线分组调度任务和无线电地图预测任务上的实验表明,与基于均值的序贯测试相比,aeSFT使用少得多的真实测试样本就能识别出有用的合成数据,其性能与固定样本符号翻转测试和配对t检验相当,同时将假阳性率保持在目标水平以下。
英文摘要
Digital twins (DTs) and learned world models are increasingly used to generate synthetic data that augment the scarce real datasets available for training artificial intelligence (AI) models in engineering systems. Owing to the inevitable simulation-to-reality (sim-to-real) gap, however, augmentation may fail to improve the performance of the trained model on the real data distribution. This paper addresses the resulting decision problem: Given a real dataset, a candidate synthetic dataset, and a fixed learning algorithm, decide whether training on the augmented dataset improves the true, population-level performance, while consuming as few real test data points as possible. Two formulations are considered: a direct test on the mean loss difference between the two trained models, and a symmetry-based test on the paired loss difference, which trades a stronger null assumption for faster evidence accumulation. For the latter, we introduce the {adaptive e-process sign-flip test} (aeSFT), a doubly adaptive procedure that adapts both the number of Monte Carlo sign-flip rounds, and hence the computational cost, and the amount of real test data consumed. aeSFT yields anytime-valid Type-I error control, with no need to pre-specify the test-set size. Experiments on a synthetic-data classification task, a DT-aided wireless packet-scheduling task, and a radio-map prediction task show that aeSFT identifies useful synthetic data using substantially fewer real test samples than mean-based sequential testing, matches the power of fixed-sample sign-flip testing and the paired $t$-test, while keeping the false-positive rate below the target level.
发表机构
- Southeast University(东南大学)
- Purple Mountain Laboratories(紫金山实验室)
- Aalborg University(奥尔堡大学)
- Northeastern University London(伦敦东北大学)
机构由 AI 辅助整理,请以论文原文为准。