协变量偏移与概念偏移的通用量化
General Quantification of Covariate and Concept Shifts
浏览论文内容
中文总结 AI 辅助
本文提出基于熵正则化最优传输的$\gamma^{*}$-概念偏移定义,统一协变量与概念偏移的误差界,并开发可估计的DataShifts算法,为分布偏移下的学习误差提供通用量化工具。
中文摘要 AI 辅助
分布偏移下的泛化仍是现代机器学习中的核心挑战,然而现有的学习界理论局限于狭窄的理想化设置,且无法从样本中估计。在本文中,我们弥合了理论与实际应用之间的鸿沟。我们首先表明,当源域和目标域的支撑集不匹配时,现有的概念偏移定义会失效。利用熵正则化最优传输,我们提出了一个关键概念:$\gamma^{*}$-概念偏移,并推导出一个统一的误差界,该误差界同时涵盖协变量偏移和$\gamma^{*}$-概念偏移,适用于广泛的损失函数、标签空间和随机标注。我们进一步为这些偏移开发了具有集中保证的估计器,以及DataShifts算法,该算法能够在大多数应用中量化分布偏移并估计误差界——这是一个用于分析分布偏移下学习误差的严谨且通用的工具。
英文摘要
Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap between theory and practical applications. We first show that existing definition of concept shift breaks when the source and target supports mismatch. Leveraging entropic optimal transport, we propose a key notion: $γ^{*}\!$-concept shifts, and derive a general error bound unifying covariate and $γ^{*}\!$-concept shifts, which applies to broad loss functions, label spaces, and stochastic labeling. We further develop estimators for these shifts with concentration guarantees, and the DataShifts algorithm, which can quantify distribution shifts and estimate the error bound in most applications - a rigorous and general tool for analyzing learning error under distribution shift.
发表机构
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。