ChartRevive:使用多模态大语言模型从图表图像重建数据可视化
ChartRevive: Reconstructing Data Visualizations from Chart Images Using MLLM
- Arizona State University(亚利桑那州立大学)
- Eli Lilly and Company(礼来公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对静态图表图像难以重用的问题,提出ChartRevive混合主动系统,利用多模态大语言模型提取数据与视觉设计,并通过交互式验证界面支持高效检查、纠正和重建图表。
AI中文摘要:
静态图表图像广泛用于科学出版物、商业报告和演示文稿中,然而从图表图像中恢复底层数据和视觉设计仍然是一个劳动密集型的手动过程,这使得它们难以被重用。虽然先前的工作主要集中在数据提取上,但视觉设计规范(包括颜色、标记形状和坐标轴配置)的提取仍未得到充分探索。为了确定适合图表重建的模型,我们系统地评估了五种多模态大语言模型(MLLMs)在五种基本图表类型上的数据和设计提取任务。我们的评估表明,文本和分类信息通常可以可靠地提取,而数值和空间信息仍然具有挑战性。在评估的模型中,GPT-5.4取得了最佳的整体性能,并被采用为我们系统的骨干。在这些发现的指导下,我们提出了ChartRevive,一个混合主动系统,它将基于MLLM的提取与交互式验证界面相结合,通过基于叠加的验证和实时重建,支持用户高效地检查、纠正和完善重建的图表。
英文摘要:
Static chart images are widely used in scientific publications, business reports, and presentations, yet recovering both the underlying data and visual design from chart images remains a labor-intensive manual process, making them difficult to reuse. While prior work has primarily focused on data extraction, the extraction of visual design specifications, including colors, marker shapes, and axis configurations, remains underexplored. To identify a suitable model for chart reconstruction, we systematically benchmark five multimodal large language models (MLLMs) across five basic chart types on both data and design extraction tasks. Our evaluation shows that textual and categorical information can generally be extracted reliably, whereas numeric and spatial information remain challenging. Among the evaluated models, GPT-5.4 achieves the best overall performance and is adopted as the backbone of our system. Guided by these findings, we present ChartRevive, a mixed-initiative system that combines MLLM-based extraction with an interactive verification interface, supporting users to efficiently inspect, correct, and refine reconstructed charts through overlay-based verification and real-time rebuilding.