arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00018cs.DBcs.AI

CITBench:面向大语言模型的交互式表格数据处理综合基准

CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs

Zihan Nan, Yang Gu, Wei Liu, Xi Yan, Zhou Liu, Hao Liang, Wentao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出 CITBench 交互式表格数据处理基准,评估发现现有 LLM 在简单表格任务表现良好,但在复杂多轮交互场景下的理解、规划与表格结构感知仍存在显著挑战。

中文摘要 AI 辅助

表格数据处理是数据工作的核心,基于大语言模型(LLM)的助手近期在支持此类任务上展现出良好能力。然而现有基准主要聚焦于单轮、完全指定指令下的表格推理,未能充分体现随用户需求演变的多轮交互复杂表格处理。为填补该缺口,本文提出 CITBench,一个用于评估 LLM 交互式表格数据处理能力的综合基准。CITBench 涵盖表格匹配、清洗、增强、转换四大类别的全面分类,包含 18 种任务类型及从不同领域数据集整理的 1296 个实例。该基准支持离线与在线评估,其中在线设置模拟受限操作流程与结构化任务脚本下的多轮交互,捕捉用户参与的表格数据处理的关键潜在行为特征。本文在 CITBench 上评估了大量开源与闭源 LLM,揭示出一致趋势:当前模型在简单表格与规则上表现良好,但随着表格复杂度提升、规则依赖增强及多轮交互模拟出现噪声,性能显著下降。这些结果凸显了 LLM 在扩展交互式数据处理场景中,在理解、规划及表格结构感知方面仍存在持续挑战。

英文摘要

Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primarily focus on table reasoning under single-turn, fully specified instructions, underrepresenting complex table processing that unfolds through multi-turn interactions with evolving user requirements. To bridge this gap, we introduce CITBench, a comprehensive benchmark for evaluating LLMs on interactive tabular data processing. CITBench features a comprehensive taxonomy across four high-level categories--table matching, cleaning, augmentation, and transformation--spanning 18 task types and 1,296 instances curated from datasets across diverse domains. The benchmark supports both offline and online evaluation, where the online setting models multi-turn interactions under constrained operation procedures and structured task scripts, capturing key potential behavioral characteristics of user-in-the-loop tabular data processing. We evaluate a broad suite of open-source and closed-source LLMs on CITBench, revealing a consistent trend: while current models perform well on simple tables and rules, their performance degrades significantly with increasing table complexity, tighter rule dependencies, and noisy multi-turn interaction simulations. These results highlight persistent challenges in understanding, planning, and table-structure awareness for LLMs in extended interactive data processing scenarios.

发表机构

  • Peking University(北京大学)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑