arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AnnoBench:可视化标注生成基准测试

AnnoBench: A Benchmark for Visualization Annotation Generation

Md Rahat-uz-Zaman, Md Dilshadur Rahman, Andrew McNutt, Paul Rosen

arXiv 2607.25911首次发表:更新:

发表机构

University of Utah(犹他大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍AnnoBench可视化标注基准测试,它将可视化与标注任务配对,通过VLM评判执行,经多因素实验评估对标注质量的影响,为推进标注自动化等提供基础。

AI 中文摘要

标注是自动化要求最高的可视化任务之一,因为它同时需要正确处理视觉、语义和风格约束。现有的基准测试或评估框架都未测试这些条件是否得到满足。我们引入了AnnoBench,这是一个可视化标注基准测试,以结构化和可测试的方式体现该领域的固有挑战。它将专业数据新闻和可视化图库中的可视化与标注任务配对,涵盖多种格式、条件和级别。通过VLM作为评判执行基准测试,并通过实验评估其对标注质量的影响。这项工作为推进标注自动化、工具和可视化生成管道提供了基础。

英文摘要

Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints. Failure to meet any of these conditions severely undermines the utility of an annotation, rendering it challenging to read, inaccurate, or visually discordant. Despite a growing body of annotation tools and automations, no existing benchmark or evaluation framework tests whether these conditions are met because of their scope and annotation not being the focus of their studies. We introduce AnnoBench, a benchmark for visualization annotation that materializes the inherent challenges of this domain in a structured and testable manner. AnnoBench pairs visualizations from professional data journalism and visualization galleries with annotation tasks, spanning four representation formats, five chart description conditions, and two prompt specification levels. The benchmark is executed via VLM-as-a-judge, using models aligned with manual human assessment. We evaluate the benchmark via four one-factor-at-a-time experiments, exploring the effects of input representation, semantic context, and prompt specificity, and model selection on annotation quality. This work provides a foundation for advancing annotation automation, tooling, and visualization-generation pipelines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑