arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LegendBench:用于图例理解与反事实干预的诊断基准

LegendBench: A Diagnostic Benchmark for Legend Understanding with Counterfactual Interventions

Xinnuo Zhang, Zhike Tang, Jing Xu, Haoyuan Zhao, Weikai Yang

arXiv 2609.24172首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLM图例理解缺乏细粒度诊断的问题,提出LegendBench基准,通过能力分类与反事实干预生成测试用例,定位并改善图例绑定与推理瓶颈。

AI 中文摘要

图例对于图表理解至关重要,因为可靠的解读需要将图例条目正确绑定到相应的视觉标记上。尽管视觉语言模型(VLM)越来越多地应用于图表理解,但其图例理解能力却难以通过总体准确率进行诊断,因为总体准确率可能被表面捷径所满足,并将图例特定错误与其他推理失败相混淆。为了实现细粒度诊断和受控测试,我们引入了LegendBench,一个参数化基准和生成流水线,用于生成针对图例的测试用例。LegendBench贡献了(1)一个能力-任务分类法,涵盖图例解析、图例定位、图例条件推理和图例感知的弃权(不执行),以定位失败;(2)反事实组生成,其中每个基础图表在受控图例干预下产生多个变体,以探测模型的不变性和敏感性。使用LegendBench,我们评估了通用VLM和专用图表模型,并生成其能力概况,揭示了在可靠的图例到标记绑定和反事实一致性方面的持续瓶颈。然后,我们利用这些能力概况来指导有针对性的微调,证明瓶颈特定的干预可以有效缩小局部能力差距并泛化到未见数据。我们进一步利用反事实设计进行细粒度诊断实验,分析编码通道效应、图例顺序捷径以及在不同可见性下的弃权(不执行)行为。

英文摘要

Legends are fundamental to chart understanding, as reliable interpretation requires correctly binding legend entries to corresponding visual marks. While vision-language models (VLMs) are increasingly applied to chart understanding, their legend understanding is poorly diagnosed by aggregate accuracy, which can be satisfied by superficial shortcuts and confound legend-specific errors with other reasoning failures. To enable fine-grained diagnosis and controlled testing, we introduce LegendBench, a parametric benchmark and generation pipeline that produces targeted legend-centric test cases. LegendBench contributes (1) a capability-task taxonomy spanning legend parsing, legend grounding, legend-conditioned reasoning, and legend-aware abstention to localize failures, and (2) counterfactual group generation, where each base chart yields multiple variants under controlled legend interventions to probe model invariance and sensitivity. Using LegendBench, we evaluate both general-purpose VLMs and specialized chart models and generate their capability profiles, revealing persistent bottlenecks in reliable legend-to-mark binding and counterfactual consistency. We then use these capability profiles to guide targeted fine-tuning, demonstrating that bottleneck-specific interventions can effectively close the localized capability gaps and generalize to unseen data. We further leverage our counterfactual design to conduct fine-grained diagnostic experiments, analyzing encoding-channel effects, legend-order shortcuts, and abstention under varying visibility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑