arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解基于大语言模型在不完美表格上的问答中的错误

Understanding Errors in LLM-Based Question Answering over Imperfect Tables

Baowen Zhang, Wei Fan, Ruman Wang, Hangting Ye

arXiv 2610.04687首次发表:更新:

发表机构

University of Wisconsin–Madison; University of Auckland; Jilin University(威斯康星大学麦迪逊分校; 奥克兰大学; 吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控实验揭示,在不完美表格问答中,错误位置影响发现率,且仅提供位置不足以保证准确;GBDI工作流结合多视图发现与验证指导,显著提升问答准确性。

AI 中文摘要

我们通过在三个人工审核的RADAR-T示例上的受控研究,调查了在不完美表格上的问答中的错误发现和处理。回答这些表格上的问题需要处理可能影响答案的错误。我们改变行顺序,并比较原始、错误标记和修复后的表格,以测试错误发现是否依赖于错误出现的位置,以及提供错误位置是否足以进行准确的问答。首先,即使表格内容和黄金答案保持不变,重新排列行也会改变错误发现。对于后置位置,完整发现率高于前置位置,并且在测试的平均位置中,紧凑布局的完整发现率高于宽间距布局。其次,仅提供经过验证的错误位置不足以进行准确的问答,在错误标记和修复后的表格之间留下了显著的准确性差距。提供已经应用了人工审核修复的表格,使得三个系统的代码辅助问答准确性比错误标记的表格提高了39.0-59.1个百分点。GBDI,一个简单的工作流程,通过将跨洗牌表格视图的错误发现与验证和处理报告错误的明确指导相结合,将这些发现付诸实践。在RADAR-T上,GBDI在五个系统上将观察到的问答准确性比代码代理基线提高了3.8-18.5个百分点。这些结果强调了在不完美表格上的问答中,可靠错误发现和有效错误处理的重要性。我们的匿名仓库可在以下网址获取:https URL

英文摘要

Answering questions over imperfect tables requires handling errors that can affect the answer. We investigate two challenges for large language models (LLMs): whether error discovery depends on where errors appear in a table, and whether providing their locations is sufficient for accurate question answering (QA). Using human-reviewed instances from RADAR-T, we conduct controlled studies across three LLMs by varying row order and comparing original, error-marked, and repaired tables. First, reordering rows changes error discovery even when the table contents and gold answer remain unchanged. During direct inspection, LLMs are more likely to discover all rows containing relevant errors when these rows appear later in the table or are grouped more closely together. Second, providing verified error locations alone is insufficient for accurate QA: with code execution, accuracy on repaired tables exceeds that on error-marked tables by 39.0-59.1 percentage points across the three LLMs. As a practical application of these findings, we combine error discovery across shuffled table views with explicit guidance for verifying and handling the reported errors in a simple workflow, Geometry-Balanced Discovery and Intervention (GBDI). On RADAR-T, GBDI improves QA accuracy by 3.8-18.5 percentage points over a code-agent baseline across five LLMs (paired 95% confidence intervals exclude zero for four), at the cost of additional inference. These results highlight the importance of both reliable error discovery and effective error handling in QA over imperfect tables. Code is available at https://github.com/645-t/GBDI-ICLR-2027.

Comments41 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑