arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RUBRIC:用于不平衡分类的现实-效用平衡排序

RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification

Yanxuan Yu, Dong Liu, Shu Wang, Wenxiao Zhao, Eric Jiang, Chang Liu, Jinxi Yu, Hui Pan, Renata Borovica-Gajic, Ben Lengerich

arXiv 2607.09816首次发表:更新:

发表机构

Columbia University; University of California, Los Angeles; University of Melbourne(哥伦比亚大学; 加利福尼亚大学洛杉矶分校; 墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究类别不平衡问题,提出RUBRIC框架,将合成样本选择设为质量优化问题,用现实-效用权衡排序,能收紧泛化边界,经实验验证其在信用卡欺诈检测等任务中可提升F1宏和召回率,保持ROC-AUC并可进行lambda敏感性分析。

AI 中文摘要

类别不平衡在欺诈检测和医疗诊断等风险敏感应用中构成了根本性挑战,其中少数类样本稀缺但对准确分类至关重要。现有过采样方法生成合成样本以重新平衡类别分布,但常产生大量低质量候选样本,导致过拟合和泛化能力下降。本文引入RUBRIC,这是一个与生成器无关的过滤框架,将合成样本选择表述为质量优于数量的优化问题。RUBRIC使用现实-效用权衡对候选样本进行排序:现实由区分真实样本和合成样本的学习鉴别器量化,效用通过基于凹边距的评分函数捕获与决策边界的接近程度。我们表明,在温和的正则条件下,所提出的过滤策略通过联合减少分布偏移和抑制近负尾贡献,单调收紧基于边距的分类器的泛化边界。通过在信用卡欺诈检测和其他不平衡基准上的广泛实验,我们证明RUBRIC提高了F1宏和召回率,同时在多个生成器上保持了可比的ROC-AUC。我们还提供了明确的lambda敏感性分析,以展示当优先考虑排序质量时用户如何恢复AUPRC。

英文摘要

Class imbalance poses a fundamental challenge in risk-sensitive applications such as fraud detection and medical diagnosis, where minority-class samples are scarce yet critical for accurate classification. Existing oversampling methods generate synthetic samples to rebalance class distributions; however, they often produce large numbers of low-quality candidates that distort decision boundaries or introduce artifacts, leading to overfitting and degraded generalization. In this work, we introduce \textbf{RUBRIC}, a generator-agnostic filtering framework that formulates synthetic sample selection as a quality-over-quantity optimization problem. RUBRIC ranks candidates using a realism-utility trade-off: realism is estimated via a neural density-ratio discriminator from each candidate's resemblance to real minority samples, while utility captures proximity to the decision boundary through a concave, margin-based scoring function $g_-(t)=-\log(1+e^{-t/τ})$. The discriminator uses the same architecture and training protocol on every benchmark and is fit independently to that dataset's real minority class versus its synthetic pool. We show that, under mild regularity conditions, the proposed filtering framework monotonically tightens the generalization bound for margin-based classifiers by jointly reducing distribution shift and suppressing near-negative tail contributions. Through extensive experiments on standard public imbalanced-classification benchmarks, we demonstrate that RUBRIC boosts minority-class recall while preserving overall discriminative ability across multiple data generators. Sensitivity analyses in $λ$ and the selection budget $K$ further characterize performance trade-offs oriented toward ranking quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑