发表机构
Carnegie Mellon University; Atlassian(卡内基梅隆大学; 阿特lassian)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对电子表格集合问答,提出先查找后计算(FiCo)方法,结合文档检索、工作簿消歧与SQL执行,在DataBench和MiMoTable上显著优于基线,并揭示源选择成本。
AI 中文摘要
电子表格集合上的问答需要找到正确的工作簿,并对完整表格进行计算。我们提出了“先查找后计算”(Find-then-Compute,FiCo)方法,该方法检索文档摘要、消除相似工作簿的歧义,并在选定的完整表格上执行受约束的结构化查询语言(SQL)。在DataBench(80个数据集,1810个问题)上,FiCo达到了76.2%的准确率:在轨道预指定的评估器下,比使用相同冻结工作簿选择的强TableRAG风格基线(66.3%)高出9.9个百分点,比前缀RAG(12.8%)高出63.4个百分点。在508个MiMoTable问题上,FiCo达到79.7%,而前缀RAG仅为22.2%。将黄金工作簿提供给强基线,使其从66.3%提升至76.3%,揭示了在固定计算下10.0个百分点的源选择成本。尽管文档召回率达到95.1%,可执行SQL达到98.5%,但只有81.3%的问题在黄金工作簿上成功执行。FiCo的优势在于将语义源选择与精确的、基于模式的計算相整合。
英文摘要
Question answering over spreadsheet collections requires finding the correct workbook and computing over complete tables. We introduce Find-then-Compute (FiCo), which retrieves document summaries, disambiguates similar workbooks, and executes constrained Structured Query Language (SQL) over the selected full table. On DataBench (80 datasets, 1,810 questions), FiCo reaches 76.2% accuracy: 9.9 points above a strong TableRAG-style baseline on the same frozen workbook choices (66.3%) under the tracks' prespecified evaluators, and 63.4 points above prefix RAG (12.8%). On 508 MiMoTable questions, FiCo reaches 79.7%, versus 22.2% for prefix RAG. Giving the strong baseline the gold workbook raises it from 66.3% to 76.3%, exposing a 10.0-point source-selection cost under fixed compute. Despite 95.1% document recall and 98.5% executable SQL, only 81.3% of questions execute on the gold workbook. FiCo's advantage comes from integrating semantic source selection with exact, schema-grounded computation.