AI 中文总结
研究表格数据缺失值预测问题,提出自动填充方法,通过后训练三个针对不同能力的专业小语言模型并结合校准集成机制,相比现有推理模型在保证高精度的同时大幅降低成本。
AI 中文摘要
预测表格数据中的缺失单元格值是数据清洗中的一个基本问题。虽然当前的推理模型在预测表格中的缺失值方面显示出很大的潜力,但通过跨行列进行整体推理,它们在大规模部署时成本高昂,并且往往过于自信,经常产生幻觉或假阳性预测。在本文中,我们观察到在表格中实现高精度的缺失值预测需要三种能力的独特组合:(1)世界知识,(2)基于文本的推理,以及(3)基于代码的推理。我们系统地探索了结合这些能力的设计选择,并提出了一种自动填充方法,该方法对三个专业的小语言模型进行后训练,每个模型针对一种能力进行优化。我们开发了一种校准集成机制,该机制可以动态选择最自信的专家或弃权,以确保高精度。在来自不同领域的2200个真实表格的11个基准上进行的广泛实验表明,与当前的推理模型(例如,o3-pro、Gemini 3 Pro和DeepSeek R1)相比,自动填充实现了更高的准确性,同时运行成本仅为这些前沿模型的一小部分(不到1%)。我们的结果突出了专业化和校准弃权在表格数据这一重要领域中的有效性。自动填充可在这个https URL上公开获取。
英文摘要
Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-art reasoning models show great promise in predicting missing values in tables, by reasoning holistically across rows and columns, they are costly to deploy at scale and tend to be overconfident, often generating hallucinated or false-positive predictions. In this paper, we observe that achieving high-precision missing-value prediction in tables requires a distinct combination of three capabilities: (1) world knowledge, (2) text-based reasoning, and (3) code-based reasoning. We systematically explore design choices for combining these capabilities, and propose an Auto-Fill approach that post-trains three specialist small language models (SLMs), each optimized for one capability. We develop a calibrated ensemble mechanism that either dynamically selects the most confident specialist or abstains, ensuring high accuracy. Extensive experiments on 11 benchmarks with 2200 real tables drawn from diverse domains show that Auto-Fill achieves superior accuracy compared to state-of-the-art reasoning models (e.g., o3-pro, Gemini 3 Pro, and DeepSeek R1), while operating at a fraction (less than 1%) of the cost of these frontier models. Our results highlight the effectiveness of specialization and calibrated abstention in the important domain of tabular data. Auto-Fill is publicly available at https://github.com/lyrain2001/auto-fill.
CommentsVLDB 2026