发表机构
National University of Defense Technology(国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出闭环先验选择框架,将LLM蒸馏因果图作为先验注入进行预算约束优化,在Do-PFN上实现2.75倍显著增益,使先验选择成为可验证问题。
AI 中文摘要
因果效应估计探究的是在干预下结果将如何变化,医学、经济学和公共政策都将其视为一项基础性任务。先验数据拟合网络(PFNs)将这一任务摊销化:一个在大量程序化生成的合成因果任务上训练的模型,将新问题的观测数据读入上下文,并在单次前向传播中返回干预效应估计。此类模型的能力很大程度上由合成训练先验决定,而该先验目前是手工设计的,这是Do-PFN和CausalPFN都承认的一个瓶颈。大型语言模型(LLMs)现在能够为给定领域“绘制”合理的因果图,这表明LLM蒸馏出的图可以作为先验材料。注入此类图是否有所帮助、收益来自何处以及何时注入有帮助,实践至今仍依赖手动试错。我们提出一个闭环先验选择框架,将先验注入视为在候选先验池上的预算约束优化。候选先验经过廉价的后训练,并由以真实领域泛化为主导的综合指标评分;胜出者随后接受完整训练并配以统计验证。在7.34M参数的Do-PFN上,该框架的胜出者在主要评估领域取得了具有统计显著性的2.75倍增益,其误差降至低于未注入的官方基线。在相邻监测领域上的泛化显著提升,且没有监测能力退化。机制实验表明,增益取决于蒸馏图的语义内容,而非仅其结构多样性(方向性证据)。借助该框架和这一规律,LLM因果先验的使用不再依赖手动试错,而成为一个可经验验证的选择问题。
英文摘要
Causal effect estimation asks how an outcome would change under an intervention, and medicine, economics, and public policy all treat it as a foundational task. Prior-data fitted networks (PFNs) amortize the task: a model trained on large numbers of programmatically generated synthetic causal tasks reads a new problem's observational data into context and returns an interventional-effect estimate in a single forward pass. The capability of such models is largely determined by the synthetic training prior, which is currently designed by hand, a bottleneck acknowledged by both Do-PFN and CausalPFN. Large language models (LLMs) can now ``draw'' plausible causal graphs for a given domain, suggesting that LLM-distilled graphs could serve as prior material. Whether injecting such graphs helps at all, where any gain comes from, and when injection helps. Practice has so far relied on manual trial and error. We propose a \emph{closed-loop prior selection framework} that casts prior injection as a budget-constrained optimization over a candidate prior pool. Candidates undergo cheap post-training and are scored by a composite metric dominated by real-domain generalization; the winner then receives full training and paired statistical validation. On a 7.34M-parameter Do-PFN, the framework's winner attains a formally significant $2.75\times$ gain on the primary evaluation domain, and its error falls below that of the uninjected official base. Generalization on an adjacent monitoring domain improves significantly, and no monitored capability degrades. Mechanism experiments show that the gain depends on the semantic content of the distilled graph rather than its structural diversity alone does not produce it (directional evidence). With this framework and this regularity in hand, the use of LLM causal priors stops being manual trial and error and becomes an empirically verifiable selection problem.