发表机构
Big Data and AI Platform Department, Tencent; Institute of Computer Science and Technology, Soochow University(腾讯大数据与人工智能平台部; 苏州大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于大语言模型的表格推理问题,提出ProgramTab框架,指导LLMs用Python代码预处理表格数据并进行关键内容提取,实验证明该框架能有效处理表格推理任务,性能优于基于LLM的基线。
AI 中文摘要
基于大语言模型(LLMs)的表格推理受到广泛关注,该任务需依据自然语言问题和结构化表格数据进行推理。然而,一系列问题制约其应用。以往方法因长文本建模困难和LLMs输入长度限制,面对大表格时性能显著下降。文本到SQL方法虽能从表格中高效提取关键信息并生成较小子表,但表格数据常缺乏必要结构和一致性,不适用于用SQL查询执行数学逻辑运算。我们提出ProgramTab框架,指导LLMs通过上下文学习用Python代码进行表格数据预处理,以及进行行列提取和SQL生成等重要内容提取。表格推理数据集实验结果表明,ProgramTab框架有效处理基于表格的推理任务,优于所有基于LLM的基线。
英文摘要
Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabular data, has gained widespread attention. However, a series of issues still constrain the application of this task. The previous approaches suffered from significant performance degradation when faced with large tables due to the difficulty of long text modeling and the limitation of input length for LLMs. The text-to-SQL approach is used to efficiently extract key information from tables and generate smaller sub-tables. However, tabular data, especially web tables, often lack the necessary structure and consistency, making them unsuitable for performing mathematical logic operations using SQL queries. We propose the ProgramTab framework, which guides LLMs employing in-context learning to perform tabular data preprocessing with Python code, as well as the momentous contents extraction with row and column extraction and SQL generation. The experiment results on table reasoning datasets demonstrate that the ProgramTab framework effectively deals with table-based reasoning tasks and outperforms all LLM-based baselines.
CommentsLarge Language Models, Table Reasoning, In-context Learning