发表机构
Arizona State University; Rensselaer Polytechnic Institute; Scale AI(亚利桑那州立大学; 伦斯勒理工学院; Scale AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对表格检索训练推理重排器的问题,提出TabRank框架,通过构建数据集并探索两种训练紧凑推理模型的变体,经压力测试,该方法在多表格检索数据集上显著提升性能,有效推广到多表格推理。
AI 中文摘要
检索相关表格以回答问题是结构化信息检索的关键任务。多阶段检索系统严重依赖重排器来优化第一阶段高效检索器生成的候选列表。神经重排器和基于大语言模型的重排方法因其语义理解和推理能力优于传统稀疏或密集检索模型而变得越来越重要。最近,配备显式思维链(CoT)推理的大型推理模型在非结构化段落检索的排名质量上有显著提升。本文提出TabRank,一个用于训练表格检索推理重排器的框架。首先给出一个包含6728条推理轨迹的综合数据集用于自然问题表格数据集上的表格重排。然后探索在这些推理轨迹上训练紧凑推理模型的两种变体:显式CoT蒸馏和在提示中让学生重排器以教师的推理轨迹为条件。在多个分布外泛化设置上进行压力测试。该方法在各种表格检索数据集上显著提高了性能,与基础模型相比,在混合问答(HybridQA)上Acc@10提高了30.5%,在SQA上提高了15.2%,在TabFact上提高了52.9%,在多表格问答基准的TATQA子集上提高了13.1%。值得注意的是,TabRank能有效推广到多表格推理。代码、数据和模型可在指定网址获取。
英文摘要
The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval systems rely heavily on rerankers to refine candidate lists produced by efficient first-stage retrievers. As a result, neural rerankers and LLM-based reranking methods have become increasingly important due to their superior capacity for semantic understanding and reasoning compared to conventional sparse or dense retrieval models. Recently, Large Reasoning Models (LRMs) equipped with explicit chain-of-thought (CoT) reasoning have shown strong improvements in ranking quality in unstructured passage retrieval. In this work, we present TabRank, a framework for training reasoning rerankers for Tabular Retrieval. We first present a comprehensive dataset of 6728 reasoning traces for tabular reranking on the Natural Questions Tables dataset. We then explore two variants of training a compact reasoning model on these reasoning traces: explicit CoT distillation and conditioning the student reranker on the teacher's reasoning trace within the prompt. We stress-test TabRank on several out-of-distribution generalization settings on diverse domains and multi-table scenarios. Our approach significantly improves performance across a variety of table retrieval datasets, increasing Acc@10 by 30.5% on HybridQA, 15.2% on SQA, 52.9% on TabFact, and 13.1% on TATQA subsets of the Multi-Table QA Benchmark compared to the base model. Notably, TabRank generalizes effectively to multi-table reasoning. Our code, data and models are available at https://github.com/AdarshSingh7647/TabRanker
Comments8 pages, 3 figures