arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36190cs.AI

FigAct:将科学图形转化为用于解释的主动画布

FigAct: Turning Scientific Figures into Active Canvases for Explanation

Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw, Jennifer Healey

首次发表
浏览论文内容

中文总结 AI 辅助

FigAct将静态科学图形转化为主动画布,通过分层搜索和视觉动作生成锚定解释,减少40倍令牌使用,使解释更清晰易理解。

中文摘要 AI 辅助

科学图形旨在通过视觉方式传达信息,然而多模态大语言模型(MLLMs)通常通过将图形的视觉内容转换回文本来解释它们。这要求读者手动将生成的解释映射回图形。受人们展示视觉信息方式的启发,我们引入了FigAct,一个通过直接作用于现有图形元素,将静态科学图形转化为基于问题的视觉呈现的框架。如同人类演示者,FigAct生成一系列简短的叙述,将每个叙述锚定在相应的视觉证据上,并应用视觉动作来引导观众的注意力。我们开发了一种分层搜索策略以实现高效的元素定位,将令牌使用量减少了约40倍。我们进一步使用三个针对性的奖励(分别用于锚定准确性、搜索效率和渲染质量)训练了FigAct-8B模型。我们还从真实世界科学论文中的图形构建了一个人工验证的基准,以评估MLLMs生成锚定视觉解释的能力。我们的结果证明了FigAct的有效性,并表明将科学图形视为演示画布能使解释更清晰、更易于理解。

英文摘要

Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40$\times$. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.

发表机构

  • Brown University(布朗大学)
  • Adobe Research(Adobe 研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑