arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04677stat.ME

带相依残差曲线的惩罚函数对函数回归中的估计与推断

Estimation and inference in penalized function-on-function regression with dependent residual curves

  • LMU Munich(慕尼黑大学)
  • Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

Fabian Scheipl

AI总结:

本文研究惩罚函数对函数回归中残差曲线相依性对估计与推断的影响,提出CL2协方差与NCV平滑选择方法,推荐NCV用于描述曲面形状、REML+CL2用于推断。

AI中文摘要:

R包refund中实现的惩罚函数对函数回归将所有曲线的所有网格值视为独立。残差曲线是平滑的,因此曲线内部的值是相依的。我们通过一个包含高斯、计数和二元响应的模拟研究,以及一个基于九个真实数据集构建的plasmode研究,考察了这种相依性对点估计和逐点置信区间的影响。在曲线内部相依的情况下,基于模型的95%区间覆盖率在0.63至0.82之间。在通常的REML拟合之后计算的、对惩罚帽子矩阵进行杠杆调整的曲线聚类三明治协方差(CL2),在模拟中将覆盖率提升至接近名义水平;在plasmode研究中,该协方差在八个数据集上略低于名义水平,在第九个数据集上系数曲面的覆盖率为0.84至0.91。Satterthwaite自由度弥补了剩余的大部分不足,但在该数据集上除外,且此时仅使用20条曲线,代价是区间变宽。然而,REML对系数曲面欠平滑:其误差是通过曲线分块邻域交叉验证(NCV)选择平滑参数的拟合误差的两到四倍,且其系数曲面的CL2区间较宽。以更平滑的NCV估计为中心的区间较窄但覆盖率不足,这是非参数回归中估计与逐点推断之间常见的张力。我们推荐使用NCV估计来描述曲面形状,并使用REML加CL2结合Satterthwaite自由度进行推断。对于处处为零的曲面,在相依误差下,这些区间在大约5%的网格点上排除零,且NCV估计比REML的估计更接近零五到七倍。对于拟合均值、截距和标量效应,REML加CL2区间接近名义水平,但二元截距和标量效应以及配准错误的计数除外。我们指出了refund的哪些默认设置得到证据支持。

英文摘要:

Penalized function-on-function regression, as implemented in the R package refund, treats all grid values of all curves as independent. Residual curves are smooth, so values within a curve are dependent. We study what this does to point estimates and pointwise confidence intervals, in a simulation with Gaussian, count and binary responses and in a plasmode study built on nine real datasets. Under within-curve dependence, model-based 95% intervals cover 0.63--0.82. A curve-clustered sandwich covariance with a leverage adjustment for the penalized hat matrix (CL2), computed after the usual REML fit, brings coverage close to nominal in the simulation; in the plasmode study it is a few points short on eight datasets and 0.84--0.91 for the coefficient surface on the ninth. Satterthwaite degrees of freedom recover most of the remaining shortfall except on that dataset, also with only 20 curves, at the cost of wider intervals. REML, however, undersmooths the coefficient surface: its error is two to four times that of a fit whose smoothing parameters are chosen by curve-blocked neighbourhood cross-validation (NCV), and its CL2 intervals for the surface are wide. Intervals centred at the smoother NCV estimate are narrow but undercover, the familiar tension between estimation and pointwise inference in nonparametric regression. We recommend the NCV estimate to describe the shape of the surface and REML + CL2 with Satterthwaite degrees of freedom for inference. For a surface that is zero everywhere, under dependent errors, these intervals exclude zero at about 5% of the grid points, and the NCV estimate is five to seven times closer to zero than REML's. For fitted means, intercepts and scalar effects, REML + CL2 intervals are close to nominal, except for binary intercepts and scalar effects and for misregistered counts. We state which defaults of refund the evidence supports.

补充信息

↑