arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Chartography:专业图表理解基准

Chartography: A Benchmark for Professional Chart Understanding

Suhaas Garre, Chris Mutty, Sushant Mehta, Edwin Chen

arXiv 2608.10677首次发表:更新:

发表机构

Surge AI(Surge AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出含100项任务的Chartography基准,其采用专业领域特定图表格式,评估显示前沿模型在该基准上表现远低于饱和水平,失败集中于视觉感知,研究还发布了相关任务、图像及评估代码。

AI 中文摘要

医学、工程、金融、制造业及科学界的专业人员常依据图表做出重大决策。现有图表基准未能充分衡量该能力:它们以条形图、折线图和饼图为主,依赖较短推理链,且已接近饱和,前沿模型已取得80-90%的分数。我们推出Chartography,这一包含100项任务的基准,配对来自专业实践的图表,采用标准图表基准极少包含的领域特定格式,问题由日常阅读此类图表的专业人员编写,并经另外三位专家独立验证。在对30种前沿模型配置的评估中(每项任务20次评分试验),最佳配置仅达到45.0%的平均pass@1;其余配置的分数在9.0-39.5%之间。失败集中在视觉感知方面:模型可能遗漏细微特征、误读稀疏标注轴上的数值、处理投影3D几何时出错,以及违反图表中编码的领域惯例。我们发布所有任务、图像、来源元数据和评估代码。

英文摘要

Professionals across medicine, engineering, finance, manufacturing, and the sciences often make consequential decisions from charts. Existing chart benchmarks do not sufficiently measure this ability: they are dominated by bar, line, and pie formats, rely on shorter reasoning chains, and are nearing saturation, with frontier models already scoring 80-90%. We introduce Chartography, a benchmark of 100 tasks that pair charts drawn from professional practice, in domain-specific formats that standard chart benchmarks rarely include, with questions written by professionals who read these charts for a living and independently verified by three additional experts. In an evaluation of 30 frontier-model configurations (20 scored trials per task), the best configuration reaches only 45.0% mean pass@1; the remainder span 9.0-39.5%. Failures concentrate in visual perception: models can miss nuanced features, misread values along sparsely labeled axes, mishandle projected 3D geometry, and violate domain conventions encoded in the chart. We release all tasks, images, provenance metadata, and evaluation code.

Comments16 pages, 5 figures, 5 tables. Accepted at the 2nd Workshop on Benchmarking Evidence-Aligned Multimodal Reasoning (BEAM 2), ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑