arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习波动:因果表格预训练的统计基础

Learning to Fluctuate: Statistical Foundations for Causal Tabular Pretraining

Zhiheng Zhang

arXiv 2609.26290首次发表:更新:

发表机构

School of Statistics and Data Science; Shanghai University of Finance and Economics(统计与数据科学学院; 上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出波动监督预训练(FSP),通过高效影响函数波动标记合成表格,证明其可消除标签模糊性并将风险降至$n^{-2}$,实验显示相比潜在监督RMSE降低69.8%。

AI 中文摘要

因果表格基础模型通过合成机制摊销效应估计,但潜在效应监督奖励后验收缩,而非直接编码固定部署总体中所需的重复样本响应。我们引入波动监督预训练(FSP):每个合成表格由其平均处理效应加上其有效影响函数波动标记,而部署仍为单次冻结前向传播。沿路径 $T_{\lambda,P}=\theta(P)+\lambda P_n\psi_P$,我们证明了一个端点转变:每个固定的 $\lambda<1$ 保留阶为 $(1-\lambda)^2/n$ 的标签模糊性,而完全波动使高斯标签可观测并将最优有限分层因果标签预测风险降至阶 $n^{-2}$。一个有限预训练界限结合了标签、网络、情节采样和优化误差;其产生的采样缺陷控制固定机制偏差、均方误差、方差、高斯近似,以及(在方差头准确时)学生化覆盖。互补下界将部署观测无法消除的局部 $n^{-1}$ ATE 风险与通用有限字典情节学习问题的 $\log N/M$ 超额风险区分开来。实验追踪了学习到的采样响应。使用原始行/列骨干,FSP 在匹配表格上将大效应偏移 RMSE 相对于潜在监督降低了 69.8%,相对于已发布的 CausalPFN 检查点降低了 39.5%。连续协变量实验、已知效应半合成和两项随机研究评估将采样律保真度与点风险收缩区分开来,并暴露了两个学习头中的弱重叠误差。

英文摘要

Causal tabular foundation models amortize effect estimation across synthetic mechanisms, but latent-effect supervision rewards posterior shrinkage rather than encoding the repeated-sample response needed in a fixed deployment population. We introduce fluctuation-supervised pretraining (FSP): each synthetic table is labeled by its average treatment effect plus its efficient influence-function fluctuation; deployment remains a frozen forward pass. Along the path $T_{λ,P}=θ(P)+λP_nψ_P$, we prove an endpoint transition: every fixed $λ<1$ retains label ambiguity of order $(1-λ)^2/n$, whereas full fluctuation makes the Gaussian label observable and reduces optimal finite-stratum causal label-prediction risk to order $n^{-2}$. A finite-pretraining bound combines label, network, episode-sampling, and optimization errors; its sampling defect controls fixed-mechanism bias, mean squared error, variance, Gaussian approximation, and, with variance-head accuracy, studentized coverage. Complementary lower bounds separate local $n^{-1}$ ATE risk from the $\log N/M$ excess risk of generic finite-dictionary episode learning. Experiments trace the learned sampling response. Across 24 nonlinear continuous-covariate cells at trained context lengths, continuous-row FSP lowers checkpoint-mean macro RMSE by 7.0% versus S-learner and wins all 12 weak-overlap cells; validation-selected Summary FSP deploys $11.6\times$ faster per table in our warm one-thread benchmark. Under effect shift, matched Raw FSP lowers mean-checkpoint RMSE by 54.2% and teacher defect by 99.0% versus latent-effect supervision, and RMSE by 10.2% versus the released CausalPFN-S checkpoint. Known-effect semisynthesis tests coverage; two randomized-study evaluations show that lower RMSE can coexist with residual attenuation.

Comments56 pages, 22 figures. Revised manuscript with expanded experiments, comparator suite, and reproducibility documentation; theoretical conclusions unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑