arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VisAudit: 评估多模态智能体在视觉诊断与修复中的表现

VisAudit: Evaluating Multimodal Agents for Visual Diagnosis and Repair

Shicheng Liu, Adam Kahirov, Qi Zhang, Zhimin Hu, Song Wang, Junhong Lin, Julian Shun, Yada Zhu

arXiv 2610.02399首次发表:更新:

AI 中文总结

针对多模态智能体在可视化自主审查中的不足,提出 VisAudit 基准,含 1,900 个缺陷实例和 300 个正确图表,评估诊断、修复和验证,实验显示最强模型仅恢复 47.4% 缺陷图表。

AI 中文摘要

多模态智能体越来越多地被用于数据可视化任务,但在自主审查方面仍然受限。与人类不同,它们可能无法识别可视化何时不正确、确定要更改的内容、在不破坏正确内容的情况下进行修复,并验证干预是否成功。现有基准主要评估预定义的单项能力,如图表生成、指令引导编辑或缺陷检测,因此未能捕捉自主审查中的这一差距。我们引入了 VisAudit,一个用于评估可视化诊断、修复和验证的基准。给定渲染的图表和可配置的辅助证据(包括其源数据表、预期文本摘要和可视化代码),智能体迭代地诊断潜在缺陷、修改并执行可视化代码、检查执行和视觉反馈,并确定何时无需进一步干预。VisAudit 定义了三个轨道,涵盖诊断修复、自主修复和开放世界验证,并包含 21 种图表类型和 10 种缺陷类别中的 1,900 个缺陷实例,以及 300 个最初正确的图表。我们通过对已验证的源可视化进行受控扰动来构建基准,并通过系统验证和人工对齐的质量控制,确保注入的缺陷定义明确且可从可用证据中恢复。与领先多模态模型的实验揭示了与可靠自主审查之间的巨大差距:最强的评估模型在自主修复设置中仅完全恢复了 47.4% 的缺陷图表。

英文摘要

Multimodal agents are increasingly used for data visualization tasks but remain limited in autonomous review. Unlike humans, they may fail to recognize when a visualization is incorrect, determine what to change, repair it without disrupting correct content, and verify whether the intervention succeeded. Existing benchmarks largely evaluate predefined individual capabilities such as chart generation, instruction-guided editing, or defect detection, and therefore do not capture this gap in autonomous review. We introduce VisAudit, a benchmark for evaluating visualization diagnosis, repair, and verification. Given a rendered chart and configurable auxiliary evidence, including its source data table, intended text summary, and visualization code, an agent iteratively diagnoses potential defects, modifies and executes visualization code, inspects execution and visual feedback, and determines when no further intervention is needed. VisAudit defines three tracks spanning diagnosed repair, autonomous repair, and open-world verification, and contains 1,900 flawed instances across 21 chart types and 10 flaw categories, together with 300 initially correct charts. We construct the benchmark through controlled perturbations of validated source visualizations, with systematic verification and human-aligned quality control to ensure that injected defects are well-defined and recoverable from the available evidence. Experiments with leading multimodal models reveal a substantial gap from reliable autonomous review: the strongest evaluated model fully recovers only $47.4\%$ of flawed charts in the autonomous-repair setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑