arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38608cs.LGstat.ML

基于学习的估计:样本选择偏差下的紧刻画

Learning-Enabled Estimation: Tight Characterizations under Sample Selection Biases

  • Yale University(耶鲁大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Vikram Kher, Jane H. Lee, Anay Mehrotra, Manolis Zampetakis

中文总结 AI 辅助

本研究刻画了样本选择偏差下回归可识别的最小条件,证明即使选择过滤器不可识别回归仍可识别,并给出有限样本保证与高效算法,适用于拍卖、劳动力市场等复杂选择场景。

中文摘要 AI 辅助

我们何时能从有偏样本中学习?我们研究回归问题,其中结果仅在通过依赖于协变量和结果本身的选择过滤器后才被观测到,这是一个普遍存在的挑战,涵盖带有患者脱落的临床试验、带有自我选择的劳动力市场以及带有策略性进入的拍卖。忽略这种选择会产生系统性的有偏结论,并带来现实世界的后果。这一挑战在计量经济学和统计学中有着悠久的历史,始于Heckman开创性的两阶段模型,随后出现了众多推广。虽然这些工作为识别提供了各种充分条件,但关于何时这种回归是可能的完整刻画仍然难以捉摸。在这项工作中,我们提供了在存在样本选择偏差时回归何时可能的刻画。我们的结果建立了在哪些选择过程的函数形式下回归仍然可能所需的最小假设,这些假设在现代设置中尤其相关,因为选择机制日益复杂且不透明。作为我们刻画的一个推论,我们表明在某些设置中,即使选择过滤器本身无法被识别,回归函数也可以被识别。这一观察已经超越了几乎所有现有方法所遵循的“先估计选择过滤器,再对回归去偏”的范式。在我们的识别条件的自然加强下,我们还建立了具有显式收敛速率的有限样本估计保证,并提供了预言机高效的算法。这为这类广泛的选择问题产生了第一种通用估计方法。最后,我们探讨了我们的结果对几个研究充分的计量经济学设置的影响,这些设置具有复杂的选择机制,如带有进入成本的拍卖和劳动力市场。

英文摘要

When can we learn from biased samples? We study regression when outcomes are observed only after passing through selection filters that depend on both covariates and outcomes themselves, a ubiquitous challenge spanning clinical trials with patient dropout, labor markets with self-selection, and auctions with strategic entry. Ignoring such selection yields systematically biased conclusions with real-world consequences. This challenge has a long history in econometrics and statistics, starting with Heckman's seminal two-stage model and followed by numerous generalizations. While these works provide various sufficient conditions for identification, a complete characterization of when such regression is possible has remained elusive. In this work, we provide a characterization for when regression is possible in the presence of sample selection bias. Our results establish the minimal assumptions required on the functional forms of selection processes under which regression remains possible, which are particularly relevant in modern settings where selection mechanisms are increasingly complex and opaque. As a corollary of our characterization, we show that there are settings where the regression function can be identified even when the selection filter itself cannot. This observation already goes beyond the ``estimate selection filter, then debias regression'' paradigm that is followed by virtually all existing approaches. Under natural strengthenings of our identification conditions, we also establish finite-sample estimation guarantees with explicit convergence rates and provide oracle-efficient algorithms. This yields the first general-purpose estimation method for this broad class of selection problems. Finally, we explore the implications of our results for several well-studied econometric settings with complex selection mechanisms such as auctions with entry costs and labor markets.

补充信息

↑