发表机构
School of Mathematical Sciences, Zhejiang University(浙江大学数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对连续输出目标偏移问题,提出RKHS谱正则化密度比估计方法,达到极小极大最优速率,并分析其在重要性加权回归中的误差传播,建立有限样本理论。
AI 中文摘要
我们研究了在连续输出目标偏移下的密度比估计和重要性加权回归问题。在目标偏移下,给定输出时输入的条件分布在训练分布和测试分布之间保持不变,而输出的边缘分布可能发生变化。尽管该问题在离散输出情形下已被广泛研究,但连续设置的理解程度明显不足:重要性权重由一个未知的密度比函数决定,而现有的估计方法缺乏显式的有限样本收敛速率。我们提出了一种在再生核希尔伯特空间(RKHS)中的谱正则化方法,用于从带标签的训练样本和无标签的测试输入中估计连续密度比。在正则性参数为$\iota>0$的源条件下,我们建立了高概率的有限样本保证,并证明该估计器达到了与容量无关的极小极大最优RKHS范数速率$O(n_\eta^{-\iota/(2\iota+2)})$。然后,我们将估计的密度比纳入重要性加权回归,并刻画了密度比估计误差向最终预测器的传播。当有足够多的样本可用于密度比估计时,所得的回归估计器达到了标准核回归的极小极大最优速率。这些结果为连续密度比估计和目标偏移下的重要性加权学习建立了有限样本理论。
英文摘要
We study density ratio estimation and importance-weighted regression under target shift with continuous outputs. Under target shift, the conditional distribution of the inputs given the outputs remains invariant across the training and test distributions, while the output marginal distribution may change. Although this problem has been extensively studied for discrete outputs, the continuous setting is substantially less understood: the importance weights are determined by an unknown density ratio function, for which existing estimation methods lack explicit finite-sample convergence rates. We propose a spectral regularization method in a reproducing kernel Hilbert space (RKHS) for estimating the continuous density ratio from labeled training samples and unlabeled test inputs. Under a source condition with regularity parameter $ι>0$, we establish high-probability finite-sample guarantees and show that the estimator achieves the capacity-independent minimax-optimal RKHS-norm rate $O(n_η^{-ι/(2ι+2)})$. We then incorporate the estimated density ratio into importance-weighted regression and characterize the propagation of density-ratio estimation error to the final predictor. When sufficiently many samples are available for density ratio estimation, the resulting regression estimator attains the minimax-optimal rates of standard kernel regression. These results establish a finite-sample theory for continuous density ratio estimation and importance-weighted learning under target shift.