arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16429cs.LG

Localized TabICLv2:通过k近邻扩展表格数据的上下文学习

Localized TabICLv2: Scaling Tabular In-Context Learning through k-NN

Beimnet Bekele Guta

首次发表
浏览论文内容

中文总结 AI 辅助

Localized TabICLv2通过仅检索k个最近训练邻居降低TabICLv2推理成本,经微调后在TabArena分类任务上保留98.64%精度,批量和单查询场景分别实现2.18倍、约249倍中位数加速。

中文摘要 AI 辅助

近年来,针对表格数据的基础模型取得了显著进展,其中TabICLv2在多项表格分类任务中达到了当前最优性能。然而,全上下文表格上下文学习(ICL)仍存在注意力成本随训练上下文规模增长的问题,这限制了其高效处理大型数据集的能力。Localized TabICLv2提出了一种方法,通过为每个测试点仅检索k个最近的训练邻居(基于模型第2阶段行表示空间的相似度进行衡量),而非使用全部训练上下文,以此降低TabICLv2的推理成本。该方法无需修改模型架构,且我们表明通过额外的第2阶段和第3阶段微调可提升精度保留效果。在TabArena分类任务上,经微调的本地化模型保留了全量TabICLv2 98.64%的精度,在批量推理中实现了中位数2.18倍的加速,在单查询服务场景中则达到了约249倍的中位数加速。

英文摘要

Foundational models for tabular data have made significant progress in recent years, with TabICLv2 reporting state-of-the-art performance on several tabular classification tasks. However, full-context tabular ICL still suffers from attention cost that grows with the training-context size, which limits its ability to handle large datasets efficiently. Localized TabICLv2 introduces a method that reduces the inference cost of TabICLv2 by retrieving only the k nearest training neighbours for each test point, measured by similarity in the model's Stage 2 row-representation space, rather than using the full training context. This requires no architectural changes, and we show that accuracy retention can be improved through additional Stage 2 and Stage 3 fine-tuning. On TabArena classification tasks, the fine-tuned localized model retains 98.64% of Full TabICLv2 accuracy and it achieves a median 2.18$\times$ speedup in batch inference, and reaches approximately 249$\times$ median speedup in the single-query serving setting.

发表机构

  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑