arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉语言模型中的图表欺骗:从漏洞到缓解

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mizanur Rahman, Mir Tafseer Nayeem, Enamul Hoque

arXiv 2607.22600首次发表:更新:

发表机构

York University; University of Alberta(约克大学; 阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型对图表欺骗的鲁棒性,引入VisDeception基准和欺骗分数,发现模型易受影响,提出推理时多智能体缓解框架,揭示可靠性差距,确立相关评估及推理方向以开发更可信的视觉分析VLM。

AI 中文摘要

信息可视化广泛用于传达模式、趋势和异常值,但诸如截断或倒置轴、扭曲的宽高比、不适当的编码和误导性颜色映射等欺骗性设计选择会系统地改变解释。随着视觉语言模型(VLM)越来越多地用于图表理解和分析推理,评估它们对这种欺骗性可视化的鲁棒性对于可信数据分析至关重要。我们引入了VisDeception,这是第一个用于评估VLM对误导性图表设计鲁棒性的受控配对基准。该基准包含1600对忠实和误导性图表,跨越八种主要的欺骗性可视化策略类别。为了将欺骗引起的推理错误与基线图表理解错误区分开来,我们引入了欺骗分数,这是一种配对评估指标。通过对10个先进VLM的32000个响应进行研究,发现即使是先进模型也极易受到欺骗性视觉操作的影响。为了提高鲁棒性,我们进一步提出了一个推理时多智能体缓解框架。我们的研究结果揭示了当前图表理解系统中重要的可靠性差距,并确立了基准驱动评估、欺骗感知指标和结构化推理作为开发更可信视觉分析VLM的有前途的方向。

英文摘要

Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematically alter interpretation while preserving the underlying data. As Vision-Language Models (VLMs) are increasingly used for chart understanding and analytical reasoning, assessing their robustness to such deceptive visualizations has become critical for trustworthy data analysis. We introduce VisDeception, the first controlled paired benchmark for evaluating the robustness of VLMs to misleading chart designs. The benchmark contains 1,600 paired faithful and misleading charts spanning eight major categories of deceptive visualization tactics, where each misleading chart is paired with a faithful counterpart generated from the same underlying data. To isolate deception-induced reasoning errors from baseline chart-understanding errors, we introduce the Deception Score, a paired evaluation metric that quantifies how misleading visualizations shift model responses away from the faithful interpretation of the data. Across 32,000 responses from 10 state-of-the-art VLMs, we find that even advanced models remain highly vulnerable to deceptive visual manipulations. To improve robustness, we further propose an inference-time multi-agent mitigation framework that grounds reasoning in structured chart metadata extracted from the visualization before answer generation, enabling models to reduce the influence of deceptive visual cues without requiring explicit user instructions. Together, our findings reveal important reliability gaps in current chart-understanding systems and establish benchmark-driven evaluation, deception-aware metrics, and structured reasoning as promising directions for developing more trustworthy VLMs for visual analytics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑