ChartJudgeBench:评估用于图表到代码生成的大型多模态模型裁判
ChartJudgeBench: Evaluating LMM Judges for Chart-to-Code Generation
浏览论文内容
中文总结 AI 辅助
ChartJudgeBench提出诊断性基准,含1,003个成对比较和650个接受/拒绝实例,评估LMM裁判在图表到代码中的可靠性,发现位置偏差、过度接受、风格匹配难及宽容偏差等局限。
中文摘要 AI 辅助
构建强大的图表到代码系统越来越依赖于强化学习,其有效性关键取决于奖励信号的质量。大型多模态模型(LMMs)在联合评估图表视觉外观和任务要求方面自然扮演着关键角色。因此,它们越来越多地被用作视觉批评者和奖励模型,但它们作为裁判的可靠性在很大程度上仍未得到探索。为此,我们引入了ChartJudgeBench,一个用于评估图表到代码工作流中LMM裁判的诊断性视觉语言基准。它包含1,003个图表感知对齐(CPA)实例,用于成对图表比较,以及650个图表推理判断(CRJ)实例,用于图表复制和图表编辑中的二元接受/拒绝验证。这些任务共同模拟了智能体优化和基于RL的图表优化中所需的核心裁判决策。我们对强LMM的评估揭示了四个系统性局限:(i)成对比较中的位置偏差,(ii)强烈倾向于过度预测接受,(iii)难以匹配视觉风格和美学,以及(iv)RL训练模型中意外的宽容偏差。这些发现表明,当前的LMM裁判在用作图表到代码优化中的批评者或奖励模型之前,需要明确的可信度验证。代码和数据可在ChartJudgeBench上获取。
英文摘要
Building strong chart-to-code systems increasingly relies on reinforcement learning, whose effectiveness depends critically on the quality of the reward signal. Large Multimodal Models (LMMs) play a natural critical role in jointly assessing chart visual appearance and task requirements. They are therefore increasingly used as visual critics and reward models, yet their reliability as judges remains largely unexplored. To this end, we introduce ChartJudgeBench, a diagnostic vision-language benchmark for assessing LMM judges in chart-to-code workflows. It includes 1,003 Chart Perception Alignment (CPA) instances for pairwise chart comparison and 650 Chart Reasoning Judgment (CRJ) instances for binary Accept/Reject verification in Chart Reproduction and Chart Editing. Together, these tasks emulate the core judging decisions required in agentic refinement and RL-based chart optimization. Our evaluation of strong LMMs reveals four systematic limitations: (i) positional bias in pairwise comparison, (ii) a strong tendency to overpredict Accept, (iii) difficulty in matching visual styles and aesthetics, and (iv) an unexpected leniency bias in RL-trained models. These findings show that current LMM judges require explicit reliability validation before being used as critics or reward models in chart-to-code optimization. The code and data are available on ChartJudgeBench.
发表机构
- Central South University(中南大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。