arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DEEPCHART:大型语言模型在忠实的数据科学图表生成方面存在多大差距?

DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen

arXiv 2608.26757首次发表:更新:

发表机构

University of Science and Technology of China; Huawei Technologies Co., Ltd.(中国科学技术大学; 华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出专家标注的DEEPCHART基准,将图表生成分为提取-推理-可视化流程,发现现有LLMs生成的图表常隐藏数据幻觉,仅靠大上下文窗口无法实现忠实图表生成,需可靠的证据提取与定量推理。

AI 中文摘要

在实际的数据科学工作流中,忠实的图表生成需要将可视化内容建立在零散证据的基础上、计算可用于生成图表的数值并准确渲染。现代大型语言模型(LLMs)能够生成视觉上合理、符合指令要求的图表,但在长文本、嘈杂且多模态的上下文环境中,数据层面的幻觉仍难以检测。为了衡量这一差距,我们推出了DEEPCHART,这是一个由专家标注的基准,包含1482个任务条件下的图表生成实例,这些实例来自实际的科学论文、金融文件和生态系统报告。DEEPCHART将图表生成表述为“提取-推理-可视化”流程,并分阶段评估源数据提取、衍生数据推理和图表渲染。对最先进模型的实验表明,视觉上合理的图表往往隐藏着数据层面的幻觉,在现实的长文本和多模态环境中,提取和推理错误十分常见。这些发现表明,仅靠更大的上下文窗口是不够的;忠实的图表生成还需要在渲染前具备可靠的证据提取和定量推理能力。我们的基准及相关资源可在此httpsURL获取。

英文摘要

Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately. Modern LLMs can produce visually plausible, instruction-compliant charts, yet data-level hallucinations remain difficult to detect in long, noisy, and multimodal contexts. To measure this gap, we introduce DEEPCHART, an expert-annotated benchmark of 1,482 task-conditioned chart-generation instances drawn from real-world scientific papers, financial filings, and ecosystem reports. DEEPCHART formulates chart generation as an Extract--Reason--Visualize pipeline and evaluates source-data extraction, derived-data reasoning, and chart rendering stage by stage. Experiments with state-of-the-art models show that visually plausible charts often conceal data-level hallucinations, with extraction and reasoning errors common in realistic long and multimodal settings. These findings suggest that larger context windows alone are insufficient; faithful chart generation also requires reliable evidence extraction and quantitative reasoning before rendering. Our benchmark and associated resources are available at https://github.com/tangdouer1005/DeepChart.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑