arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23926stat.MLcs.LG

密度比重新评分用于不平衡分类

Density-Ratio Rescoring for Imbalanced Classification Using Raking Duals and Classifier Scores

  • Sungshin Women’s University(诚信女子大学)
  • Kangwon National University(江原国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Dongha Kim, Seunghwan Park

AI总结:

本研究提出密度比重新评分(DRR)方法,利用加权对偶分数增强原始分类器,在不重采样下提升不平衡分类的稀有类排序,在24个基准上平均精度增益0.034。

AI中文摘要:

密度比重新评分(DRR)通过在原始类别先验下训练的分类器上增加一个基于调查加权(survey-raking)的对偶分数来增强分类器。加权法在容差范围内重新加权多数类样本,以匹配少数类的特征矩。DRR 对偶分数和基础分数进行边缘标准化,并以固定权重二分之一进行组合,直接使用拟合的对偶分数进行预测,无需重采样或重新拟合基础分类器。在精确总体匹配和正确指定的对数线性倾斜模型下,对偶分数等于对数密度比加上一个加性常数。类别分离分析刻画了在常见类内协方差下,融合提升分离度的信号强度和相关性条件。在24个表格基准上,经过30次试验和五个基础学习器评估,DRR在D=128随机特征设置下,在每个数据集上相对于标准化基础分数提升了平均精度,平均增益为0.034。它在24个数据集中的22个上超过了共享对偶加权和重标记重采样器,平均增益为0.092,并在共享基因表达队列的所有八个一对多任务中均超过。这些结果证明了使用加权对偶分数作为可重用分数来改进稀有类别排序的有效性,同时保留了在原始先验下训练的分类器。

英文摘要:

Density-Ratio Rescoring (DRR) augments a classifier trained at the original class prior with a survey-raking dual score. Raking reweights the majority sample to match minority feature moments within a tolerance. DRR marginally standardizes the dual and base scores and combines them with a fixed weight of one half, using the fitted dual directly for prediction without resampling or refitting the base classifier. Under exact population matching and a correctly specified log-linear tilt model, the dual equals the log density ratio up to an additive constant. A class-separation analysis characterizes the signal strength and correlation conditions under which fusion improves separation under common within-class covariance. On 24 tabular benchmarks, evaluated over 30 trials and five base learners, DRR at the D=128 random-feature setting improves average precision over the standardized base on every dataset, with a mean gain of 0.034. It exceeds the shared-dual raking-and-relabeling resampler on 22 of 24 datasets, with a mean gain of $0.092$, and on all eight one-versus-rest tasks of a shared gene-expression cohort. These results demonstrate the effectiveness of using raking duals as reusable scores for improving rare-class ranking while retaining classifiers trained at the original prior.

补充信息

↑