arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时调查感知交叉验证很重要?是组内相关系数,而非设计效应

When Does Survey-Aware Cross-Validation Matter? The ICC, Not the Design Effect

M. Ehsan Karim, Md Belal Hossain

arXiv 2607.10634首次发表:更新:

AI 中文总结

研究复杂样本调查中交叉验证问题,通过结果和初步线性预测器的组内类内相关系数诊断,在三项全国健康调查中评估朴素与尊重设计的交叉验证配对,证明诊断有效,指出分层方法问题及权重处理错误。

AI 中文摘要

K折交叉验证假定观测值可交换,但复杂样本调查的分层、聚类和不等权重违反了这一假设。存在尊重设计的“调查CV”,但额外关注何时会改变任何结论的问题仍未解决。我们在验证任何模型之前,通过两种低成本诊断方法来回答这个问题:结果和初步线性预测器的组内类内相关系数(ICC)。一个模拟的阳性对照证明了它们的敏感性,随着ICC的增加,朴素交叉验证对新聚类性能变得乐观,而聚类级别的折保持真实。然后,我们在三项全国健康调查中评估了朴素交叉验证与尊重设计的交叉验证的配对情况(慢性疼痛、糖尿病和青少年自杀倾向预测;惩罚和随机森林学习器,加上一个无惩罚比较器)——一次重新分析、一次前瞻性应用和一次预先指定的筛选。在所有三项中,诊断方法都正确地预测了结果:没有实际大小的方案差异,唯一排除零的区间显示出悲观,而不是聚类泄漏产生的乐观——即使设计效应很大(与ICC不同,这不是正确的触发因素)。我们还表明,分层方法在公共使用设计中通常不可行,并给出了一个备用层次结构,我们记录了权重处理错误,其数量级的伪影使任何折方案效应相形见绌。本文附有可重现代码。

英文摘要

K-fold cross-validation assumes exchangeable observations, violated by the stratification, clustering, and unequal weights of complex sample surveys. Design-respecting "survey CV" exists, but the question of when the extra care changes any conclusion has remained open. We answer it before validating any model, with two inexpensive diagnostics: the within-cluster intraclass correlation (ICC) of the outcome and of a preliminary linear predictor. A simulated positive control demonstrates their sensitivity, with naive cross-validation growing optimistic about new-cluster performance as the ICC rises while cluster-level folds stay honest. We then evaluate paired naive-versus-design-respecting cross-validation in three national health surveys (chronic-pain, diabetes, and adolescent-suicidality prediction; penalized and random-forest learners, plus an unpenalized comparator) - one reanalysis, one prospective application, and one prespecified screen. In all three the diagnostics correctly anticipated the outcome: no scheme difference of practical size, and the only interval excluding zero showed pessimism, not the optimism that cluster leakage produces - even where the design effect was large (which, unlike the ICC, is not the right trigger). We also show the stratified recipe is often infeasible in public-use designs and give a fallback hierarchy, and we document weight-handling errors whose order-of-magnitude artifacts dwarfed any fold-scheme effect. Reproducible code accompanies the paper.

Comments9 pages, 2 figures, 1 table. Supporting Information (Web Appendices A-D) is provided as an ancillary file

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑