arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于非结构化数据的预训练嵌入计量经济学

Econometrics with Pre-Trained Embeddings for Unstructured Data

Yuya Shimizu

arXiv 2607.17378首次发表:更新:

发表机构

University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究非结构化数据在实证经济学中的应用,针对使用预训练模型提取嵌入作为协变量存在的理论基础有限问题,给出克服困难的充分条件及收敛速度,通过双重机器学习应用估计相关参数。

AI 中文摘要

非结构化数据,如图像和文本,在实证经济学中使用得越来越多。由于在非结构化数据上训练机器学习模型成本高昂,经济学家常使用计算机科学家开发的现成预训练深度学习模型来提取嵌入,然后将其用作目标经济分析中的协变量。尽管这种做法很流行,但其理论基础仍然有限。存在两个主要困难。首先,预训练模型通常在不同数据集上针对不同任务进行训练,因此不清楚何时该模型可可靠用于目标任务。其次,嵌入函数存在识别问题,这使得难以分析嵌入函数的估计误差及其对目标任务的影响。在本文中,我们提供了克服这些困难的充分条件,并推导了具有预训练嵌入的机器学习模型的收敛速度。我们通过双重机器学习应用来说明该理论,这些应用用于估计感兴趣的参数,如具有非结构化控制的部分线性回归、考虑由图像和文本测量的产品质量的需求估计中的价格弹性、使用非结构化数据的缺失数据插补以及具有非结构化混杂因素的平均处理效应。

英文摘要

Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machine-learning models on unstructured data is costly, economists often use off-the-shelf pre-trained deep learning models developed by computer scientists to extract embeddings, which are then used as covariates in target economic analyses. Despite the popularity of this practice, its theoretical foundations remain limited. There are two main difficulties. First, pre-trained models are typically trained on different datasets and for different tasks, making it unclear when they can be used reliably for the target task. Second, the embedding function is subject to an identification problem, complicating the analysis of its estimation error and the effect of that error on the target task. We provide sufficient conditions to overcome these difficulties. A key condition, which we call transferability, governs the convergence rate we derive. To assess transferability, we develop a computationally feasible bootstrap test that does not require re-estimating the embeddings and nuisance functions. Our theory applies to a wide range of double machine learning applications, including partially linear regression with unstructured controls, price elasticity estimation in demand models accounting for product quality measured by images and text, missing-data imputation using unstructured data, and average treatment effect estimation with unstructured confounders. As an empirical application, we estimate the labor supply elasticity on Amazon Mechanical Turk, an online labor market platform, using job-description embeddings as controls.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑