arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25789stat.MEmath.STstat.TH

分组半监督估计两样本泛函:基于单指标条件插补

Grouped Semi-Supervised Estimation of Two-Sample Functionals via Single-Index Conditional Imputation

  • College of Mathematics and Systems Science, Xinjiang University(新疆大学数学与系统科学学院)

机构由 AI 辅助整理,请以论文原文为准。

Hengyao Xu, Tao Tan

AI总结:

针对异质分组中两样本比较泛函估计,提出结合单指标条件插补与协变量平均的半监督方法,在理论上保证一致性与渐近正态性,并通过模拟和NHANES数据验证其偏差小、均方误差低及标准误更小的优势。

AI中文摘要:

在异质分组中估计两样本比较泛函具有挑战性,因为结果仅在一小部分标记子集中被观测到,而协变量广泛可得。我们开发了分组特定两样本比较泛函及其预设聚合的半监督估计。所提出的估计器将来自标记交叉臂对的单指标条件插补与对所有可用协变量对的平均相结合。为了处理可能在可达端点处消失的指数得分密度,我们使用中心化稳定化,在剖面准则中使用平滑稳定器,并在点估计器中对核分母使用收缩的下截断(硬下限)。外部平均保留所有协变量对。在正确指定的单指标条件均值模型和所述正则条件下,我们建立了有界比较泛函的一致性和渐近正态性。一个四角色原始观测影响表示解释了样本重叠,并识别了未标记协变量相对于监督估计可以减少的方差分量。我们进一步建立了具有完全重拟合和共享标记计数的观测级多项扰动重抽样的条件有效性,区分了理想方差一致性与有限重复方差估计。模拟显示,在点估计设置中,偏差较小且均方误差低于监督估计,同时扰动实验中的覆盖率接近名义水平。一个描述性的NHANES 2011-2018示例,比较不同年龄组中按高血压病史分层的血清肌酐,其点估计与监督对应方法相似,且所有报告比较的半监督标准误更小。

英文摘要:

Estimating two-sample comparison functionals within heterogeneous groups is challenging when outcomes are observed only for a small labelled subset while covariates are widely available. We develop semi-supervised estimation of group-specific two-sample comparison functionals and their prespecified aggregates. The proposed estimator combines single-index conditional imputation from labelled cross-arm pairs with averaging over all available covariate pairs. To handle index-score densities that may vanish at attainable endpoints, we use centred stabilization, with a smooth stabilizer in the profile criterion and a shrinking lower truncation of the kernel denominator (the hard floor) in the point estimator. The outer average retains all covariate pairs. Under a correctly specified single-index conditional-mean model and the stated regularity conditions, we establish consistency and asymptotic normality for bounded comparison functionals. A four-role original-observation influence representation accounts for sample overlap and identifies the variance component that unlabelled covariates can reduce relative to supervised estimation. We further establish conditional validity of observation-level multinomial perturbation resampling with complete refitting and shared labelled counts, distinguishing ideal variance consistency from finite-replicate variance estimation. Simulations show small bias and lower mean squared error than supervised estimation across the point-estimation settings, together with near-nominal coverage in the perturbation experiment. A descriptive NHANES 2011-2018 illustration comparing serum creatinine by hypertension history across age groups yields point estimates similar to their supervised counterparts and smaller semi-supervised standard errors for all reported comparisons.

补充信息

↑