ARASH:面向表格预测的自适应检索与样本选择
ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction
浏览论文内容
中文总结 AI 辅助
本文提出ARASH方法,通过训练集局部邻域分析选择最优样本,在保持TabPFN相当准确率的同时,大幅降低其提示长度与内存使用。
中文摘要 AI 辅助
表格预测是众多应用中的关键任务,大型语言模型的近期成功催生了多种将其适配到表格领域的方法。一种流行策略是训练或微调专用表格基础模型(Tabular Foundation Models, TFMs),例如TabPFN。但TFMs需要大量计算资源,频繁重新训练往往不切实际。上下文学习(In-Context Learning, ICL),尤其是少样本提示,是提升性能的资源高效替代方案。然而,针对表格数据,确定最相关的行作为样本(shots)仍是一项挑战。本文提出ARASH(Adaptive, query-specific Retrieval And Shot selection,自适应查询特定的检索与样本选择),该方法通过基于训练集内的局部邻域分析选择最优样本,提升TFMs的效率。实验结果表明,ARASH将TabPFN的提示长度和内存使用分别降低1261.5倍和2.56倍,同时提供相当的准确率。
英文摘要
Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approaches for adapting them to the tabular domain. A prevalent strategy involves training or fine-tuning specialized Tabular Foundation Models (TFMs) such as TabPFN. However, TFMs require substantial computational resources, and frequent model retraining is often impractical. In-context learning (ICL), specifically, few-shot prompting, offers a resource-efficient alternative to enhance performance. Yet, identifying the most relevant rows to serve as shots remains a challenge for tabular data. This paper introduces ARASH (Adaptive, query-specific Retrieval And Shot selection), a method that improves TFM efficiency by selecting optimal shots based on local neighborhood analysis within the training set. Our results demonstrate that ARASH reduces the prompt length and memory usage of TabPFN by 1261.5$\times$ and 2.56$\times$, respectively, while providing comparable accuracy.
发表机构
- McMaster University(麦克马斯特大学)
机构由 AI 辅助整理,请以论文原文为准。