发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出LLM数据分析工具存在忽视数据探索步骤的差距,引入两个基准评估数据集理解,发现强模型仍会遗漏逻辑结构,显式数据探索可提升下游任务性能。
AI 中文摘要
基于大语言模型(LLM)的数据分析工具正越来越多地用于帮助用户分析杂乱的电子表格和工作簿,范围从对上传文件的问题回答到生成代码、摘要和可视化。这些系统通常通过最终下游答案的正确性来评估。然而,可靠的数据分析还依赖于更早的步骤:在解决请求的任务之前理解数据集包含的内容。对于复杂工作簿,此数据探索步骤包括识别物理工作表背后的逻辑表、解释列语义、恢复键和关系以及检测质量问题。在当前的工具和基准中,此步骤通常被隐含处理,在下游任务性能与可靠、可人工检查分析所需的数据集理解之间造成了差距。我们的主要贡献是识别这一被忽视的差距,将数据探索确立为一级评估目标,并通过下游实验表明,更强的数据探索支持可提升任务性能。为直接评估数据集理解,我们引入了两个基准设置:基于维生素D研究数据集的真实多工作表工作簿基准,以及带有固定模式数据探索工件的DSBench扩展。在这两个设置中,系统通过捕获表、列、语义角色、关系和分析信号的结构化工件的质量进行评估。我们的结果表明,即使读取电子表格内容,强大的LLM和数据分析智能体仍会遗漏重要的逻辑结构。此外,显式的数据探索支持通常会提升下游正确性,这表明应将其视为LLM数据分析工作流中的一级可检查阶段,以及自然的人工介入检查点,领域专家可在下游分析进行前审查并修正该工件。
英文摘要
LLM-based data-analysis tools are increasingly used to help users analyze messy spreadsheets and workbooks, from answering questions over uploaded files to generating code, summaries, and visualizations. These systems are often evaluated by the correctness of their final downstream answers. However, reliable data analysis also depends on an earlier step: understanding what the dataset contains before solving the requested task. For complex workbooks, this Data Exploration step includes identifying the logical tables behind physical sheets, interpreting column semantics, recovering keys and relationships, and detecting quality issues. In current tools and benchmarks, this step is usually left implicit, creating a gap between downstream task performance and the dataset understanding needed for reliable, human-checkable analysis. Our key contribution is to identify this overlooked gap, make Data Exploration a first-class evaluation target, and show through downstream experiments that stronger Data Exploration support improves task performance. To evaluate dataset understanding directly, we introduce two benchmark settings: a real multi-sheet workbook benchmark based on a Vitamin D study dataset, and an extension of DSBench with schema-fixed Data Exploration artifacts. In both settings, systems are evaluated by the quality of a structured artifact capturing tables, columns, semantic roles, relationships, and profiling signals. Our results show that strong LLMs and data-analysis agents still miss important logical structure even when they read spreadsheet content. Furthermore, explicit Data Exploration support often improves downstream correctness, suggesting it should be treated as a first-class, inspectable stage in LLM data-analysis workflows and a natural human-in-the-loop checkpoint where domain experts can review and correct the artifact before downstream analysis proceeds.
Comments9 pages, 6 figures. Accepted to VLDB 2026 Workshop: DASHSys: Systems for Data-centric Agents with Human-in-the-loop