arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02909econ.EMstat.ML

当预测成为回归元:针对下游推断偏差的拆分样本校正方法

When Predictions Become Regressors: A Split-Sample Correction for Biases in Downstream Inference

Nathan Canen, Ted Enamorado

AI总结:

本文针对预测生成测度用作回归元的偏差问题,提出基于独立拆分样本构建工具变量的校正方法,经模拟及两个实证案例验证其有效性。

AI中文摘要:

基于预测的方法,包括大型语言模型(Large Language Models, LLMs)及其他机器学习技术,常被用于构建难以直接量化的政治现象测度,如竞选纲领中的政策立场、社交媒体上表达的情绪等。在诸多应用场景中,这些由预测生成的测度会被用作回归模型的解释变量,即便其存在测量误差,这会导致估计结果出现偏差。本文提出一种针对此类偏差的简单解决方案:基于原始数据的独立拆分样本所生成的多个测度构建工具变量。该方法在理论上成立、易于实现且无需新数据。通过模拟实验,我们表明该方法即使在样本量相对较小时,也能恢复出接近真实值的估计结果,而标准方法在实际应用中会产生显著偏差。我们通过两个应用案例说明该方法:一是德国议会中性别化言论是否影响立法结果,二是政治风险是否影响中国的减贫项目。

英文摘要:

Prediction-based methods, including Large Language Models (LLMs) and other machine learning techniques, are often used to construct measures of political phenomena that are difficult to quantify directly, such as policy positions in manifestos or emotions expressed on social media. In many applications, these prediction-generated measures are used as explanatory variables in regression models, even though they are measured with error. This leads to biased estimates. In this paper, we propose a simple solution to these biases: instrumental variables constructed from multiple measures created on independent splits of the original data. This approach is theoretically valid, easy to implement, and does not require new data. Through simulations, we show that this approach recovers estimates close to the true values, even in relatively small samples, while the standard approach can produce substantial bias in practice. We illustrate the method by revisiting two applications: whether gendered speech affects legislative outcomes in the German Parliament, and whether political risk influences poverty alleviation programs in China.

↑