发表机构
American Institutes for Research; Northwestern University(美国研究院; 西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究潜在回归变量经多个非线性测量时的线性回归问题,通过固定潜在尺度、限制曲率异质性确定结构系数区间,可通过辅助回归估计限制,以美国社区调查面板数据为例应用,得到相关系数及半宽度,视为测量协调。
AI 中文摘要
我们研究当回归变量是潜在的且仅通过多个噪声测量来观察时的线性回归,每个测量都是潜在变量的平滑但可能非线性的函数。在人工智能职业暴露测量中该问题很严重,不同分数导致下游估计相差11倍。对单个测量进行回归得到的是特定源系数而非结构系数。我们通过要求共识测量函数为线性来固定潜在尺度,并限制各源相对于斜率的剩余曲率异质性。在此限制下,结构系数位于以对称跨源估计器为中心的封闭形式区间内。该区间对未知源负荷不变,其半宽度在曲率限制中是二阶的且精确到相同阶数。至少有四个测量时,可通过分裂工具辅助回归从源的联合分布中估计该限制,使用Stoye临界值的Imbens-Manski置信区间在曲率类上实现均匀覆盖。应用将六种暴露测量与2015年至2024年的888万人年观察的美国社区调查面板相匹配。2022年后语言模型测量和Webb专利文本测量的就业系数符号不同,事前因子分析规则将Webb测量分离为一个独特的结构。五个保留源产生的负荷不变共识系数为-0.239,部分识别半宽度为点估计的1.23%,在曲率的单侧95%上限处为1.88%。我们将该应用视为测量协调而非人工智能替代的因果估计。
英文摘要
We study linear regression when the regressor is latent and observed only through multiple noisy measurements, each a smooth but possibly nonlinear function of the latent variable. The problem is acute in the measurement of occupational exposure to artificial intelligence, where competing scores yield downstream estimates that differ by a factor of eleven. A regression on any single measurement recovers a source-specific coefficient rather than the structural one. We fix the latent scale by requiring the consensus measurement function to be linear and bound the remaining curvature heterogeneity across sources relative to slope. Under this bound, the structural coefficient lies in a closed-form interval centered at a symmetric cross-source estimator. The interval is invariant to unknown source loadings, and its half-width is second order in the curvature bound and sharp to the same order. With at least four measurements, the bound is estimable from the joint distribution of the sources through a split-instrument auxiliary regression, and Imbens-Manski confidence intervals with the Stoye critical value attain uniform coverage over the curvature class, including at the point-identified boundary. The application matches six exposure measures to an American Community Survey panel of 8.88 million person-year observations for 2015 to 2024. The post-2022 employment coefficient changes sign between the language-model measures and the Webb patent-text measure, and an ex ante factor-analytic rule separates the Webb measure as a distinct construct. The five retained sources yield a loading-invariant consensus coefficient of -0.239, with a partial-identification half-width of 1.23 percent of the point estimate, or 1.88 percent at the one-sided 95 percent upper bound on the curvature. We read the application as measurement reconciliation rather than as a causal estimate of AI displacement.