发表机构
Hologen AI; Mathematics Research Centre Academy of Athens; FAU Erlangen Nuremberg; Imperial College London; University of Edinburgh(全息人工智能公司; 雅典科学院数学研究中心; 埃尔朗根 - 纽伦堡大学; 伦敦帝国理工学院; 爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究表格上下文学习器中虚假变量内信号路由问题,证明岭回归上下文学习器下该路由不可避免,推导出相关特征。发现更大上下文会放大虚假路由,更具表现力模型更脆弱。引入两种轻量级缓解措施,有效减少虚假路由并提高因果敏感性。
AI 中文摘要
考虑一个在单一医院训练以预测患者康复的模型,其中测量特征\(X\)将患者的真实健康信号\(C\)与该医院设备的系统伪影\(S\)捆绑在一起。在该医院内,伪影通过未测量的混杂因素(如患者人口统计学)与结果相关;上下文学习器会合理地通过\(S\)而非\(C\)进行预测,并且在部署到具有不同设备的新医院时会悄然失败。我们将此形式化为“复合表示中的虚假路由”:当特征\(X = [C;\,\alpha S;\,\eta]\)在不同子空间中编码因果信号\(C\)和虚假信号\(S\)时,上下文学习器无法确定哪个驱动预测。我们证明,在岭回归上下文学习器(一种线性上下文学习器)下,无论上下文大小如何,这种路由都是不可避免的;最先进的预训练表格上下文学习器TabPFN在经验上显示出定性一致的行为。我们推导出一个封闭形式的特征,\(\mathrm{CSR} \propto \rho_S/\rho_C\),线性上下文学习器在\(r = 0.997\)时得到证实,TabPFN在\(r = 0.979\)时得到证实。与直觉相反,更大的上下文会强化对主导上下文信号的依赖,将虚假路由放大高达\(1.74\times\);在高虚假角落,更具表现力的模型在经验上显示出更大的脆弱性(在高纠缠时CSR差距为\(+2.22\))。我们引入了两种轻量级缓解措施:环境分层上下文构建和\(S\)交换增强,它们只需要弱环境标签,不需要因果分区的知识。\(S\)交换将线性上下文学习器的虚假路由减少了\(74\%\),将TabPFN的虚假路由减少了\(98.8\%\),同时TabPFN的因果敏感性提高了\(8.4\times\):模型不会变得不可知,而是通过因果信号重新路由。
英文摘要
Consider a model trained at a single hospital to predict patient recovery, where the measured feature $X$ bundles the patient's true health signal ($C$) with a systematic artefact from that hospital's equipment ($S$). Within that hospital, the artefact correlates with outcomes through unmeasured confounders such as patient demographics; an in-context learner rationally routes predictions through $S$, not $C$, and fails silently when deployed at a new hospital with different equipment. We formalise this as \emph{spurious routing in composite representations}: when a feature $X = [C;\,αS;\,η]$ encodes a causal signal $C$ and a spurious signal $S$ in distinct subspaces, the ICL cannot determine which drives predictions. We prove that under ridge ICL, a linear in-context learner, this routing is unavoidable regardless of context size; TabPFN, a state-of-the-art pretrained tabular ICL model, shows qualitatively consistent behaviour empirically. We derive a closed-form characterisation, $\mathrm{CSR} \propto ρ_S/ρ_C$, confirmed at $r = 0.997$ for linear ICL and $r = 0.979$ for TabPFN. Contrary to intuition, larger context sharpens commitment to the dominant in-context signal, amplifying spurious routing by up to $1.74\times$; in the high-spurious corner, more expressive models show greater vulnerability empirically ($+2.22$ CSR gap at high entanglement). We introduce two lightweight mitigations: environment-stratified context construction and S-swap augmentation, that require only weak environment labels and no knowledge of the causal partition. S-swap reduces spurious routing by $74\%$ for linear ICL and $98.8\%$ for TabPFN, with TabPFN's causal sensitivity increasing $8.4\times$ simultaneously: the model does not become agnostic, it reroutes through the causal signal.