CITBench:面向大语言模型的交互式表格数据处理综合基准
CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs
浏览论文内容
中文总结 AI 辅助
本文提出 CITBench 交互式表格数据处理基准,评估发现现有 LLM 在简单表格任务表现良好,但在复杂多轮交互场景下的理解、规划与表格结构感知仍存在显著挑战。
中文摘要 AI 辅助
表格数据处理是数据工作的核心,基于大语言模型(LLM)的助手近期在支持此类任务上展现出良好能力。然而现有基准主要聚焦于单轮、完全指定指令下的表格推理,未能充分体现随用户需求演变的多轮交互复杂表格处理。为填补该缺口,本文提出 CITBench,一个用于评估 LLM 交互式表格数据处理能力的综合基准。CITBench 涵盖表格匹配、清洗、增强、转换四大类别的全面分类,包含 18 种任务类型及从不同领域数据集整理的 1296 个实例。该基准支持离线与在线评估,其中在线设置模拟受限操作流程与结构化任务脚本下的多轮交互,捕捉用户参与的表格数据处理的关键潜在行为特征。本文在 CITBench 上评估了大量开源与闭源 LLM,揭示出一致趋势:当前模型在简单表格与规则上表现良好,但随着表格复杂度提升、规则依赖增强及多轮交互模拟出现噪声,性能显著下降。这些结果凸显了 LLM 在扩展交互式数据处理场景中,在理解、规划及表格结构感知方面仍存在持续挑战。
英文摘要
Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primarily focus on table reasoning under single-turn, fully specified instructions, underrepresenting complex table processing that unfolds through multi-turn interactions with evolving user requirements. To bridge this gap, we introduce CITBench, a comprehensive benchmark for evaluating LLMs on interactive tabular data processing. CITBench features a comprehensive taxonomy across four high-level categories--table matching, cleaning, augmentation, and transformation--spanning 18 task types and 1,296 instances curated from datasets across diverse domains. The benchmark supports both offline and online evaluation, where the online setting models multi-turn interactions under constrained operation procedures and structured task scripts, capturing key potential behavioral characteristics of user-in-the-loop tabular data processing. We evaluate a broad suite of open-source and closed-source LLMs on CITBench, revealing a consistent trend: while current models perform well on simple tables and rules, their performance degrades significantly with increasing table complexity, tighter rule dependencies, and noisy multi-turn interaction simulations. These results highlight persistent challenges in understanding, planning, and table-structure awareness for LLMs in extended interactive data processing scenarios.
发表机构
- Peking University(北京大学)
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构由 AI 辅助整理,请以论文原文为准。