arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越表格数据的TabPFN:多模态嵌入的校准与准确性

TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings

Jingxiang Zhang, Lujia Zhong, Zijie Zhu, Shuo Huang, Yuang Xu

arXiv 2607.11007首次发表:更新:

AI 中文总结

研究少样本多模态分类中轻量级头部校准不佳问题,核心方法是评估TabPFN作为即插即用分类头。贡献在于其在多数据集、多编码器和多模态评估中表现优异,能提升校准并保持准确率,为校准敏感多模态分类提供指导。

AI 中文摘要

少样本多模态分类通常将轻量级头部(如k近邻、逻辑回归或线性支持向量机)附加到冻结的预训练编码器上。尽管计算效率高,但这些头部可能产生校准不佳的置信度分数,限制了它们在校准敏感应用中的可靠性。我们评估TabPFN作为冻结图像、文本和音频编码器的即插即用、零梯度分类头部。在跨越14个数据集、11个编码器和三种模态的22820个评估情节中,TabPFN在负对数似然(NLL)和预期校准误差(ECE)方面在九个分类头部中实现了最佳平均排名。在一个代表性设置中,相对于八个基线的平均值,它将NLL降低了48 - 62%,将ECE降低了2.1 - 5.3倍,同时匹配或超过了它们的平均准确率。其准确性优势是有条件的,集中在中等到高样本数和低到中等特征维度(k≥50,d≤32),当标记数据稀缺、特征维度高或竞争方法接近上限准确率时会减弱。在有针对性的骨干适应实验中,用TabPFN替换训练好的线性头部在保持竞争准确率的同时显著提高了校准。这些结果为在校准敏感的多模态分类中使用TabPFN作为无训练头部提供了实证指导。为了支持透明度和可重复性,我们在GitHub仓库中公开发布了源代码、实验配置和评估脚本。

英文摘要

Few-shot multimodal classification commonly attaches a lightweight head, such as $k$-nearest neighbors, logistic regression, or a linear SVM, to a frozen pretrained encoder. Although computationally efficient, these heads can produce poorly calibrated confidence scores. We ask whether TabPFN can provide reliable confidence estimates on multimodal embeddings without sacrificing predictive accuracy, and under what conditions. We systematically evaluate TabPFN as a zero-gradient head for frozen image, text, and audio encoders. Across 22{,}820 evaluation episodes spanning 14 datasets, 11 encoders, and three modalities, TabPFN achieves the best mean rank among nine classification heads on both negative log-likelihood (NLL) and expected calibration error (ECE). At a representative setting, it reduces NLL by 48--62\% and ECE by 2.1--5.3$\times$ relative to the average of eight baselines while matching or exceeding their average accuracy. This calibration benefit transfers broadly, whereas the accuracy advantage is conditional: it concentrates at moderate-to-high shot counts and low-to-moderate feature dimensions ($k \ge 50$, $d \le 32$), and diminishes when labeled data are scarce, feature dimensions are high, or competing methods approach ceiling accuracy. After backbone adaptation, replacing the trained linear head with TabPFN improves calibration while preserving competitive accuracy, showing that representation adaptation and reliable head choice are complementary. Together, these results identify when TabPFN can serve as a training-free head for calibration-sensitive multimodal classification. To support transparency and reproducibility, we publicly release the source code, experiment configurations, and evaluation scripts in our GitHub repository: https://github.com/Jingxiang-Zhang/tabpfn-multimodal-embeddings.

Comments19 pages, 13 figures, 10 tables. Jingxiang Zhang and Lujia Zhong contributed equally. Code: https://github.com/Jingxiang-Zhang/tabpfn-multimodal-embeddings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑