InfoEdit:探究信息图编辑中的全局布局推理
InfoEdit: Probing Global Layout Reasoning in Infographic Editing
AI总结:
本文提出InfoEdit基准,含1000个信息图与4000条指令,评估八种编辑器,发现仅GPT-Image-2成功率超60%,揭示回流是结构化内容编辑的核心挑战。
AI中文摘要:
多模态基础模型能够以生产级质量编辑自然照片,但同样的模型在处理诸如信息图等结构化视觉内容时却表现不佳。与照片不同,信息图通过逻辑关系编码信息;编辑一个元素往往需要调整周围元素。我们将这种全局布局推理能力称为“回流”(reflow)。现有的图像编辑基准既未为结构化视觉内容提供专门设置,也未评估回流能力。我们引入了InfoEdit,这是一个新颖的基准,包含跨越八个逻辑关系家族的1,000个信息图,配以跨越四个编辑任务的4,000条编辑指令,以及一个回流感知的评估协议。在八个前沿编辑器中,仅GPT-Image-2的平均成功率超过60%;大多数模型低于7%,且即使目标定位完美,也没有编辑器在交换块(Swap-Block)任务上超过36%。我们进一步表明,代码级编辑能够匹敌最强的像素级编辑器,揭示了任务间的互补优势。InfoEdit将回流识别为结构化视觉内容编辑中的核心挑战,并提供了一个诊断性基准以促进未来的进展。
英文摘要:
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to this global layout reasoning capability as reflow. Existing image-editing benchmarks neither provide a dedicated setting for structured visual content nor evaluate the reflow capability. We introduce InfoEdit, a novel benchmark of 1,000 infographics across eight logical-relation families, paired with 4,000 editing instructions across four editing tasks, and a reflow-aware evaluation protocol. Across eight frontier editors, only GPT-Image-2 clears 60% average success rate; most models fall below 7%, and no editor exceeds 36% on the Swap-Block task even with perfect target localization. We further show that code-level editing can match the strongest pixel-level editor, revealing complementary strengths across tasks. InfoEdit identifies reflow as a central challenge in structured visual content editing and provides a diagnostic benchmark to facilitate future progress.