表格基础模型在上下文中计算什么?通过注意力门控更新的原位表示精炼
What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates
查看机构详情
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出原位表示精炼方法,通过注意力门控更新构建任务特定预测器,RefineICL模型在多个基准上超越现有方法,并验证了支持表示更新的有效性。
中文摘要 AI 辅助
当每个表格都定义了一个新的监督任务时,表格基础模型应该学习哪种可复用的计算?我们提出了原位表示精炼(in-situ representation refinement):支持标签引导对当前情节(episode)表示的更新,这些更新无需改变模型参数即可迁移到未标记的查询上。一个正则化的留一法(leave-one-out)目标函数产生了一个支持校正项及其查询扩展项。主导项将基于注意力的读取与状态相关的缩放分离开来,这启发了RefineICL:一个注意力门控、无前馈网络(FFN-free)的上下文堆栈,具有选定的低秩特征交互和类型化记忆。RefineICL-L24在AMLB29上达到了0.93836的OVR-AUC和0.87173的准确率。在38数据集TabArena快照上,一个基准信息引导的延续训练达到了1644.8 Elo,在相同评估下比TabPFN-3高出31.4 Elo。在TabZilla的两个视图上,它还在所有四个报告的指标上优于TabPFN-v3。在一个匹配的100K更新深度网格中,扩展的FFN没有带来一致的验证收益,并且在L8时使用了60.2%更多的峰值推理内存。内部干预表明,支持表示不仅仅是标签的静态来源:移除一次中间支持更新,同时保留查询输出,在所有72个测试的情节中增加了最终查询的交叉熵。综合来看,推导过程和干预实验解释了注意力门控更新如何在上下文中构建一个任务特定的预测器。
英文摘要
A tabular foundation model must discover which distinctions matter for each new table without updating its parameters. We develop in-situ representation refinement: support labels guide changes to the episode's representations, improving the information available to later queries. A regularized leave-one-out objective yields a support correction and its query extension. The leading term separates attention-based reading from state-dependent scaling, motivating RefineICL: an attention-gated, FFN-free contextual stack with selected low-rank feature interaction and typed memory. A direct intervention tests the role of evolving support states: removing one intermediate support update while preserving the block's query output increases final query cross-entropy in all 72 tested episodes. RefineICL-L24 reaches 0.93836 OVR-AUC and 0.87173 accuracy on AMLB29. A benchmark-informed continuation reaches 1644.8 Elo on the 38-dataset TabArena snapshot, 31.4 Elo above TabPFN-3 under the same evaluation. It also improves all four reported metrics over TabPFN-v3 on both TabZilla views. In a matched 100K-update depth grid, an expanded FFN gives no consistent validation benefit and uses 60.2% more peak inference memory at L8. These results connect learning within a forward pass to representation refinement and show how this view guides a competitive, memory-efficient model.