AI 中文总结
研究针对表格问答中现有方法的局限,提出技能增强表格图推理框架,将表格表示为属性图,利用大语言模型规划执行动态链检索证据子图,构建技能库提炼抽象技能进行对比增强推理,实现持续自我进化,提升了表格问答性能。
AI 中文摘要
表格问答旨在通过表格推理回答用户查询。现有研究统一处理所有问题并仅通过整体准确率评估,掩盖了大语言模型在简单查找方面表现出色但在聚合和算术等复杂操作上存在困难的现实。为揭示这种差异,我们引入了具有细粒度问题分类法的新颖操作级表格问答任务,并发布了两个数据集WikiTQ - ow和TabFact - ow用于评估。针对现有方法将表格扁平化破坏固有结构及从头推理忽视可复用模式的问题,我们提出技能增强表格图推理框架。该框架将表格表示为具有明确行列单元格结构的属性图,大语言模型规划并执行动态链以检索证据子图进行图遍历推理,构建分层技能库提炼推理轨迹为抽象技能,进行对比增强表格图推理实现持续自我进化。实验表明该框架性能优越,平均整体提升5.91%、操作级提升6.03%,还减少了19.76%的令牌消耗和27.64%的推理延迟。
英文摘要
Table Question Answering (TableQA) aims to reason over tables to answer user queries. Existing research treats all questions uniformly and evaluates solely through overall accuracy, obscuring a critical reality that LLMs excel at simple lookups yet struggle with complex operations like aggregation and arithmetic. To reveal this disparity, we introduce a novel \emph{Operation-wise TableQA} task with a fine-grained question taxonomy and release two datasets named WikiTQ-ow and TabFact-ow for evaluation. As for modeling bottlenecks, existing methods flatten tables into linearized texts, disrupting inherent structures and inducing the ``lost-in-the-middle'' issue, which poses a primary barrier to complex cross-row reasoning. Moreover, they typically reason from scratch, neglecting reusable patterns shared across similar operations. To address these limitations, we propose a Skill-augmented Table Graph Reasoning (SkillTGR) framework for self-evolving structured reasoning. Specifically, SkillTGR represents tables as attributed graphs with explicit row-column-cell structures, where LLMs plan and execute dynamic chains to retrieve evidence subgraphs for graph traversal reasoning. Based on this, SkillTGR builds a hierarchical SkillBank to distill reason trajectories into abstract skills under cognitive heuristics, then hybrid retrieves both successful and failed skills for contrastive augmented table graph reasoning, thereby enabling the continual self-evolution. Extensive experiments demonstrate that SkillTGR achieves superior performance with an average of 5.91\% overall and 5.06\% operation-wise improvement, also reducing 19.76\% token consumption and 27.64\% inference latency. Our codes and data will be released upon publication.
Comments14 pages, 7 figures, Accepted by ICDE 2027