发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对文本因果推断中因处理状态由文本编码引发的混淆陷阱,提出掩码调整表征方法,经模拟验证可改善重叠诊断、稳定因果效应估计并降低偏差。
AI 中文摘要
从观测文本估计语言属性的因果效应颇具挑战,因为同一文档既包含目标处理变量,也包含调整所需的非处理文本属性。现有方法通常从全文学习表征以捕捉潜在混淆,但当处理状态本身由文本中的词语编码时,这些表征会直接编码处理变量,从而形成混淆陷阱:更丰富的表征可使处理组与对照组文档可分,即便基础因果问题满足重叠假设,也会引发重叠违反。本文研究通过词典或其他处理定义词汇信息编码的潜在文本处理,提出基于掩码的调整表征,在表征学习前移除此类词汇处理信号。我们形式化了表征诱导的重叠失效,证明删除掩码可保留词袋/主题模型表征的重叠,并将替换掩码表征为大型语言模型的自然松弛,其隐藏处理定义 token 的同时保留词序与上下文。在模拟实验中,与从未掩码文本学习的调整方法相比,掩码可改善重叠诊断、稳定处理效应估计并降低偏差。
英文摘要
Estimating causal effects of linguistic properties from observational text is difficult because the same document can contain both the treatment of interest and the non-treatment textual attributes needed for adjustment. Existing approaches often learn representations from the full text to capture latent confounding, but when treatment status is itself encoded by words in the text, these representations can directly encode treatment. This creates a confounder trap: richer representations can make treated and control documents separable, inducing overlap violations even when the underlying causal problem satisfies overlap. We study latent text treatments that are encoded through lexicons or other treatment-defining lexical information, and propose masking-based adjustment representations that remove this lexical treatment signal before representation learning. We formalize representation-induced overlap failure, prove that deletion masking preserves overlap for bag-of-words/topic-model representations, and characterize replacement masking as a natural relaxation for large language models that hides treatment-defining tokens while preserving word order and context. Across simulations, masking improves overlap diagnostics, stabilizes treatment effect estimates, and reduces bias relative to adjustment methods that learn from the unmasked text.