arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越参与者水平交叉验证:纵向机器学习的可靠推断

Beyond Participant-Level Cross-Validation: Reliable Inference for Longitudinal Machine Learning

Shahran Rahman Alve

arXiv 2608.26205首次发表:更新:

AI 中文总结

该研究针对纵向机器学习中伪复制等问题,提出与分析匹配的排列检验,经模拟和公开队列验证,可有效控制I类错误,明确纵向研究的分析规则。

AI 中文摘要

纵向感知研究通常从数十名参与者中收集数千个窗口,记录数量庞大,但独立科学单元并非如此。当结果按参与者定义时,这种不匹配会导致看似精确的结果易受伪复制、划分选择以及在报告某一流程前比较多个流程的常规分析灵活性的影响。按参与者划分可防止个人记录跨越划分,但无法校准产生报告数量的依赖标签的工作流折叠构建、预处理、调优、校准和候选选择。我们定义了一个参与者水平的估计量,并通过对参与者标签进行排列并重新运行整个工作流来获得与分析匹配的原假设。在受控模拟中,当不存在效应时,窗口水平推断在70%-80%的重复中拒绝,窗口自助法也以相同的速率拒绝;参与者自助法仍以10%-17%的速率拒绝;与分析匹配的检验在20至80名参与者的队列中保持0.025-0.100的水平。冻结所选流程而非重复搜索会使I类错误在有8个候选时膨胀至0.240,而重复搜索时则保持0.040。应用于两个公开队列,腕部活动记录仪(n=55)得出参与者AUROC为0.928,p=0.0050,该结论在规模稳健的秩合并统计量以及排除住院参与者后计算的匹配排列原假设下仍成立(p=0.0089)。智能手机感知(n=38,7个阳性样本)得出0.636,尽管分辨率足够,但未拒绝(p=0.1724),灵敏度为0.143。实际规则很明确:每个读取标签的步骤都应包含在排列分析中,重复记录不会产生额外的独立参与者。

英文摘要

Longitudinal sensing studies routinely collect thousands of windows from a few dozen participants. The records are numerous; the independent scientific units are not. When the outcome is defined per participant, this mismatch makes apparently precise findings vulnerable to pseudo-replication, to partition choice, and to the ordinary analytic flexibility of comparing several pipelines before reporting one. Splitting on participants prevents a person's records from straddling a split, but it does not calibrate the label-dependent workflow fold construction, preprocessing, tuning, calibration, and candidate selection that produced the reported number. We define a participant-level estimand and obtain an analysis-matched null by permuting participant labels and rerunning that entire workflow. In controlled simulation, window-level inference rejects in 70-80% of replicates when no effect exists and a window bootstrap rejects at the same rate; a participant bootstrap still rejects at 10-17%; the analysis-matched test holds 0.025-0.100 across cohorts of 20 to 80 participants. Freezing the selected pipeline instead of repeating the search inflates Type-I error to 0.240 with eight candidates, where repeating it holds 0.040. Applied to two public cohorts, wrist actigraphy (n=55) yields participant AUROC 0.928 with p=0.0050, a conclusion that persists under a scale-robust rank-pooled statistic and under a matched permutation null computed after excluding hospitalized participants (p=0.0089). Smartphone sensing (n=38, 7 positives) yields 0.636 and does not reject (p=0.1724) despite sufficient resolution, with sensitivity 0.143. The practical rule is narrow: every step that reads labels belongs inside the permuted analysis, and repeated records do not create additional independent participants.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑