发表机构
Federal Reserve Bank of Philadelphia(费城联邦储备银行)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探讨文本嵌入用于实证分析的假设,在文档为潜在主题混合的生成模型下明确该假设,通过对363个美国大都市区的应用,发现基于LLM生成经济描述的嵌入聚类效果优于其他聚类方式。
AI 中文摘要
文本嵌入何时可作为实证分析的输入?其应用基于一个假设:我们可以用文本的低维嵌入替代文本,且不会损失太多信息。我在文档为潜在主题混合的生成模型下,对该假设进行了精确表述。我研究了两种应用:在嵌入空间中对单元进行聚类,以及对高维文本进行控制。嵌入的聚类是具有相似主题混合的一组文档;控制嵌入等价于控制主题混合,因此有效性取决于该混合是否能捕捉混淆因素。在对363个美国大都市区的应用中,基于LLM生成的经济描述的嵌入聚类,可恢复可解释的经济原型,且比基于模型残差或精心挑选的行业与人口统计协变量的聚类,能更清晰地分离本地就业动态。
英文摘要
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the confounding. In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.