ChartAnno:评估多模态大语言模型的图表标注生成能力
ChartAnno: Benchmarking Multimodal Large Language Models for Chart Annotation Generation
浏览论文内容
中文总结 AI 辅助
本文提出ChartAnno基准评估MLLMs的图表标注生成能力,实验显示专有模型表现更优,大规模开源模型正缩小差距,更具体指令可提升标注质量,图表图像仅在设计相关指标上有有限提升。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在图表理解、生成与编辑领域已取得显著进展,但其对现有图表进行标注的能力却未得到充分探索。图表标注是一项常见且具有挑战性的交流任务,要求模型推断图表的预期信息、解读图表语义,并放置合适的文本或图形元素。为填补这一空白,本文提出ChartAnno——一个用于评估MLLMs图表标注生成能力的基准,它包含1200张真实世界图表,配有对应代码及三个指令特异性等级的标注指令。我们在两种主要输入设置下评估10个代表性MLLMs:(1)仅图表代码;(2)图表代码与图表图像,还额外开展了仅图表图像的 ablation 研究。结果显示,专有模型整体表现仍更优,不过大规模开源模型正缩小差距;更具体的指令可提升标注质量,而推断抽象意图仍是当前MLLMs面临的最大挑战;提供图表图像带来的整体提升有限,改进主要体现在与设计相关的指标上。这些发现表明图表标注生成是一项需要语义接地和有效标注设计的挑战性任务,代码与数据将在后续版本发布。
英文摘要
Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large language models (MLLMs) offer new opportunities for automatic chart annotation authoring, their capabilities in this task remain underexplored. To address this gap, we introduce ChartAnno, a comprehensive benchmark for evaluating MLLMs on chart annotation generation. ChartAnno contains 1,200 real-world charts with paired annotated and unannotated executable code, along with 3,600 annotation instructions spanning three levels of specificity. We also develop a multidimensional evaluation framework combining rule-based and LLM-judged metrics to assess execution, structural compliance, semantic consistency, and design effectiveness. We evaluate 10 representative MLLMs under two primary chart input settings: (1) chart code alone and (2) both code and chart image. Results reveal that proprietary models lead overall, though open-source models narrow the gap. While higher instruction specificity improves annotation quality, inferring abstract communicative intent remains difficult across all models. Providing chart images yields marginal benefit when code is available. We also examine the effect of chart code through an image-only ablation and analyze the effects of multiple task complexity indicators and instruction-level transitions. Further analyses characterize common failure modes and validate the reliability of the LLM-based judge. Experiments with D3 and SVG demonstrate the generalizability of ChartAnno beyond its primary Python setting.