arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20787cs.LGcs.AI

合成少数群体数据是冗余的或无效的:一种数据依赖的有效性理论和去偏检验

Synthetic minority data is redundant or invalid: a data-dependent validity theory and a de-biased test

  • Faculty of Information Technology, Mutah University(穆塔大学信息技术学院)
  • Department of Accounting, Mutah University(穆塔大学会计系)

机构由 AI 辅助整理,请以论文原文为准。

Ahmad B. Hassanat, Ahmad S. Tarawneh, Ghada A. Altarawneh

AI总结:

研究类别不平衡学习中合成少数群体数据的有效性,提出去偏检验方法,证明有效性取决于数据,通过实验表明多数方法不满足标准,发布检验工具并颠倒举证责任,要求合成数据证明有效性和信息增益。

AI中文摘要:

二十年来,处理类别不平衡学习的标准方法是生成合成少数群体样本,其有效性的标准证据是一种不会失败的检验:用生成合成点的数据对其进行评分。我们对该检验进行去偏。有效性成为一个总体数量——合成点真正属于少数群体类别的概率,通过用保留的真实数据对合成点进行评分的一致估计量来衡量。在有可用的保留真实情况时,经典检验在96 - 99%的方法与不平衡率单元格中低估了真正的无效性,而去偏估计量能紧密跟踪它。我们证明有效性是数据的属性,而非方法的属性:类别重叠设定了一个任何可靠生成器都无法逃脱的无效性下限,使得在类别分离时过采样冗余,在类别重叠时无效。在91种方法、三个分类器以及涵盖医学和金融的数据集上——包括一个设计用于通过经典检验的生成器——没有一个能同时满足两个标准:相对于最佳平凡基线的增益很微弱(中位数低于0.01 F1,超出决策阈值范围),且大多损害了校准。我们将该检验作为可通过pip安装的测试发布,并颠倒了举证责任:合成少数群体数据现在必须在手头的数据上证明有效性和信息增益。

英文摘要:

For two decades, the standard remedy for class-imbalanced learning has been to fabricate synthetic minority examples, and the standard evidence of their validity has been a check that cannot fail: synthetic points are scored against the very data that generated them. We de-bias the check. Validity becomes a population quantity -- the probability that a synthetic point truly belongs to the minority class -- with a consistent estimator that scores synthetic points against withheld real data. Where held-out ground truth is available, the classical test underestimates true invalidity in 96-99% of method-by-imbalance-ratio cells, while the de-biased estimator tracks it closely. We prove validity is a property of the data, not the method: class overlap sets an invalidity floor no faithful generator escapes, making oversampling redundant where classes separate and invalid where they overlap. Across 91 methods, three classifiers, and datasets spanning medicine and finance -- including a generator engineered to pass the classical check -- none clears both bars: gains over the best trivial baseline are noise-thin (median below 0.01 F1, a decision threshold's reach), and most damage calibration. We release the audit as a pip-installable test and flip the burden of proof: synthetic minority data must now demonstrate, on the data at hand, both validity and information gain.

补充信息

↑