发表机构
Center for Computational Science and Engineering, MIT; Department of Civil and Environmental Engineering, MIT; Institute for Data, Systems, and Society, MIT(麻省理工学院计算科学与工程中心; 麻省理工学院土木与环境工程系; 麻省理工学院数据、系统与社会研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文刻画了部分线性模型中系数估计的尖锐结构无关极小极大风险,通过新颖下界证明标准双重机器学习可能高估估计难度,并提出联合平衡两个干扰学习器的近似与随机误差原则。
AI 中文摘要
我们刻画了在结果和治疗干扰函数由两个不同的黑箱学习器学习时,部分线性模型中系数估计的尖锐结构无关极小极大风险,这解决了Gu (2025)在双重机器学习中提出的开放问题。对于每个干扰函数\\(q\in\{\mu,\pi\}\\),我们通过近似误差预算\\(a_q\\)和随机误差预算\\(s_q\\)来刻画可用学习器,其中后者通过局部Rademacher复杂度控制。记\\(\mathcal E_n\\)为极小极大均方误差,我们证明\\[\mathcal E_n\asymp1\wedge\left\{\frac1n+\left(a_\mu a_\pi+\min\left\{a_\pi s_\mu+s_\pi^2,\\,a_\mu s_\pi+s_\mu^2\right\}\right)^2\right\}.\\]主要的新成分是针对一般双学习器问题的一个新颖下界。我们的证明使用正交编码函数构造了四个有限混合检验实验。在这些实验中,隐藏扰动分别被放置在两个学习器类之外、仅治疗学习器类之外、仅结果学习器类之外,或两个学习器类之内。这四种配置分别捕捉了两个近似误差之间的交互、一个学习器的近似误差与另一个学习器的学习误差之间的两种不对称交互,以及同时学习两个干扰函数的联合估计难度。将四个所得下界组合起来得到显示的速率,该速率与Gu (2026)的最新上界匹配。我们的结果表明,标准双重机器学习可能高估目标估计的内在难度,并为学习器选择提供了目标特定的原则:近似误差和随机复杂度必须在两个干扰学习器之间联合平衡,而不是分别优化。
英文摘要
We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in double machine learning posed by Gu (2025). For each nuisance \(q\in\{μ,π\}\), we characterize the available learner by an approximation-error budget \(a_q\) and a stochastic-error budget \(s_q\), with the latter controlled through localized Rademacher complexity. Writing \(\mathcal E_n\) for the minimax mean-squared error, we show that \[\mathcal E_n\asymp1\wedge\left\{\frac1n+\left(a_μa_π+\min\left\{a_πs_μ+s_π^2,\,a_μs_π+s_μ^2\right\}\right)^2\right\}.\] The main new ingredient is a novel lower bound for the general two-learner problem. Our proof constructs four finite-mixture testing experiments using orthogonal code functions. Across these experiments, the hidden perturbations are placed outside both learner classes, outside only the treatment learner class, outside only the outcome learner class, or inside both learner classes. These four configurations capture, respectively, the interaction between the two approximation errors, the two asymmetric interactions between one learner's approximation error and the other learner's learning error, and the joint estimation difficulty of learning both nuisances. Combining the four resulting lower bounds yields the displayed rate, which matches the latest upper bound in Gu (2026). Our result shows that standard double machine learning can overstate the intrinsic difficulty of target estimation and provides a target-specific principle for learner selection: approximation error and stochastic complexity must be jointly balanced across the two nuisance learners rather than optimized separately.