发表机构
Renmin University of China; National University of Singapore; Rutgers University(中国人民大学; 新加坡国立大学; 罗格斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过对比不同架构的表格基础模型在无关特征抑制上的表现,发现交替轴架构更优,并证明保留可寻址特征轴为任务自适应相关性推断提供归纳偏置,支持架构与先验对齐。
AI 中文摘要
表格基础模型(TFMs)因能通过上下文学习在新数据集上提供强预测,无需任务特定训练或大量调参而日益流行。然而,已发布的TFMs在预训练先验、架构和目标上同时存在差异,掩盖了它们各自的归纳偏置。因此,我们考察一个具体能力:无关特征抑制。在合成任务和真实世界数据集上,添加空特征导致行令牌模型TabDPT的预测性能显著下降,而单元格令牌交替轴模型TabPFN v2及其他TFMs则相对稳定。这一差距促使我们探究架构是否有助于无关特征抑制。由于已发布的TFMs受其他设计选择混淆,我们在相同的稀疏到密集线性先验下训练精简的行令牌和交替轴变压器。精确贝叶斯分析表明,稀疏预测需要上下文相关的特征门控,而密集端点仅需均匀特征加权。与此区分一致,交替轴模型在稀疏任务上更接近贝叶斯最优预测器,而架构差距在密集任务上变得可忽略;稀疏差距几乎全部源于线性系数估计误差。最后,在受控模型和冻结的TabPFN v2中,我们考察了特征注意力输出对线性系数的干预效果,发现存在任务依赖的选择性计算路由通过特征索引路径的证据。这些结果共同支持架构-先验对齐:保留可寻址的特征轴为任务自适应相关性推断提供了归纳偏置。代码可在该https URL获取。
英文摘要
Tabular foundation models (TFMs) are increasingly popular because they deliver strong predictions on new datasets through in-context learning, without task-specific training or extensive tuning. Yet released TFMs differ simultaneously in their pretraining priors, architectures, and objectives, obscuring their respective inductive biases. We therefore examine one concrete capability: irrelevant-feature suppression. Across synthetic tasks and real-world datasets, adding null features causes substantially greater predictive degradation in the row-token model TabDPT, whereas the cell-token alternating-axis model TabPFN v2 and other TFMs remain comparatively stable. This gap motivates us to ask whether architecture contributes to irrelevant-feature suppression. Because released TFMs remain confounded by other design choices, we train streamlined row-token and alternating-axis transformers under identical sparse-to-dense linear priors. Exact Bayes analysis shows that sparse prediction requires context-dependent feature gating, whereas the dense endpoint requires only uniform feature weighting. Consistent with this distinction, the alternating-axis model is substantially closer to the Bayesian optimal predictor on sparse tasks, while the architecture gap becomes negligible on dense tasks; almost all of the sparse gap arises from linear coefficient-estimation error. Finally, in both the controlled model and frozen TabPFN v2, we examine the effect of interventions on the feature-attention outputs on the linear coefficients, finding evidence of task-dependent selective routing of computation through feature-indexed pathways. Together, these results support architecture-prior alignment: preserving an addressable feature axis provides an inductive bias for task-adaptive relevance inference. Code is available at https://github.com/Tianqi-Zhao/ArchitecturePriorTFMs.