arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分位数感知扩散建模对高度不平衡表格数据的有用性

Usefulness of Quantile-Aware Diffusion Modeling for Highly Imbalanced Tabular Data

Abu Talha, Peng Liu, Souradyuti Paul

arXiv 2610.05825首次发表:更新:

发表机构

Indian Institute of Technology Bhilai; Loughborough University(印度理工学院比莱分校; 拉夫堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高度不平衡表格数据,提出分位数正则化去噪的Quantile-TabDDPM方法,结合二次误差与分位数损失,在信用卡欺诈检测中表现稳健有效。

AI 中文摘要

在高度不平衡数据的背景下,分类问题是许多实际应用(如金融科技、医疗保健等)中的重大挑战。在这些情况下,绝大多数实例属于单一类别,而一小部分代表少数类别(通常是最关键的类别)。近年来,扩散模型已成为降低数据集中“不平衡程度”的强大方法;它们通过迭代变换捕获复杂的数据分布来生成合成数据。然而,标准的扩散模型本质上不适合高度偏斜或重尾的数据,由于其内置的二次误差损失,缺乏结构敏感性来捕获罕见的极端值和少数类别的细微差别。我们提出了一种新颖的方法,即Quantile-TabDDPM,基于分位数正则化的去噪目标,该目标将标准二次误差损失与分位数损失项相结合,以明确捕获罕见事件,同时保留原始去噪目标的理论基础。我们在一个以极端类别不平衡为特征的真实世界信用卡交易数据集上广泛评估了我们的方法。结果表明,基于扩散的合成数据生成与分位数正则化去噪目标的集成为高度不平衡数据集中的欺诈检测提供了一个稳健且有效的框架。

英文摘要

Classification problem in the context of highly imbalanced data is a major challenge in many real-world applications (e.g., FinTech, healthcare, etc.). In these cases, the vast majority of instances belong to a single class and a small fraction represent the minority class (often the most critical class). Recently, diffusion models have emerged as powerful approaches to reduce the degree of ``imbalanced-ness'' in the dataset; they work by generating synthetic data by capturing complex data distributions using iterative transformations. However, standard diffusion models are not inherently suited to highly skewed or heavy-tailed data, due to inbuilt quadratic error loss, which lacks the structural sensitivity to capture rare, extreme values, and minority-class nuances. We propose a novel approach, namely, Quantile-TabDDPM, based on a quantile-regularized denoising objective that combines the standard quadratic error loss with a quantile loss term to explicitly capture rare events while preserving the theoretical grounding of the original denoising objective. We extensively evaluated our approach on a real-world credit card transaction dataset characterized by extreme class imbalance. The results demonstrate that the integration of diffusion-based synthetic data generation with a quantile-regularized denoising objective provides a robust and effective framework for fraud detection in highly imbalanced datasets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑