发表机构
Islamic Azad University(伊斯兰阿扎德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出BAPS框架,无需修改预训练TabPFN模型,可将百万级表格数据集的上下文压缩约1953倍,用512个原型保留预测性能,实现其在大规模表格数据上的可扩展推理。
AI 中文摘要
预训练表格基础模型已展现出强大的预测能力,但它们在大规模数据集上的应用仍受限于有限的推理上下文。本文提出平衡自适应原型选择(Balanced Adaptive Prototype Selection, BAPS)框架,用于构建紧凑且信息保留的上下文以实现可扩展的TabPFN推理。该框架无需修改或重新训练预训练模型,可同时保留代表性结构、信息决策边界、局部密度、类别平衡和特征空间多样性。在百万行规模的HIGGS和SUSY数据集上的实验表明,512个原型可保留强大的预测性能和可靠的校准,对应约1953倍的上下文压缩。所有实验均在配备16GB RAM且无GPU加速的Intel Core i7 CPU上完成。这些发现表明,有效的上下文构建是将预训练表格基础模型扩展至百万级数据集的实用机制。
英文摘要
Pretrained tabular foundation models have demonstrated strong predictive capability; however, their application to large-scale datasets remains constrained by the limited inference context. This paper introduces Balanced Adaptive Prototype Selection (BAPS), a framework for constructing compact, information-preserving contexts for scalable TabPFN inference. Without modifying or retraining the pretrained model, BAPS jointly preserves representative structure, informative decision boundaries, local density, class balance, and feature-space diversity. Experiments on the million-row HIGGS and SUSY datasets show that 512 prototypes retain strong predictive performance and reliable calibration, corresponding to an approximately 1,953-fold context compression. All experiments were conducted on an Intel Core i7 CPU with 16 GB RAM and no GPU acceleration. These findings establish effective context construction as a practical mechanism for extending pretrained tabular foundation models to million-scale datasets.