发表机构
SZU; SIAT, CAS; UCAS; China Tower Corporation; SUAT(深圳大学; 中国科学院深圳先进技术研究院; 中国科学院大学; 中国铁塔股份有限公司; 深圳技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态编码智能体评估的缺陷,提出ReFigBench基准,用1000张arXiv真实概览图测试科学图形重建为可编辑PPT的任务,发现感知瓶颈及模型与执行框架的耦合影响,揭示保真度与可编辑性的核心矛盾。
AI 中文摘要
多模态编码智能体被期望将视觉输入转化为可用的工件,它们通过一个控制框架(即围绕模型的工具、上下文管理和执行环境的层次)来行动。现有评估往往孤立地考察短工具调用、API轨迹或截图相似度,在这些代理指标下的低分无法说明模型是观察不佳、规划不佳,还是其控制框架导致失败。我们研究科学概览图形重建,这是一个智能体任务,其中源图像必须变成保留文本、拓扑、布局和原生文档结构的可编辑PowerPoint幻灯片。我们引入ReFigBench,一个基于从arXiv论文中检索到的1,000张具有完整来源的真实概览图形构建的基准和评估框架。来自四个模型家族的编码智能体在两种工作流程(直接代码生成和专门的PPTX工作流程)下重建每张图形,最强的模型在两个商业控制框架内运行,产生十种配置。评估结合了确定性工件检查、来自两个模型家族的评判者进行的重复自动评分,以及盲法人工比较。感知仍然是一个瓶颈,迭代渲染只能部分弥补。工作流程的努力是否能转化为质量取决于模型及其控制框架,因为同一模型在一个控制框架内从专门工作流程中获益,而在另一个控制框架内则受损,并且即使在相同的直接提示下,控制框架也会改变分数。专门的流程在每种配置中都消除了原生连接符,人工评判者在大多数对决中仍然更喜欢其渲染效果,即使是最强的智能体也达不到评分标准的上限。这些结果揭示了保真度与可编辑性之间的张力,这是实用多模态文档智能体的核心挑战。
英文摘要
Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often isolate short tool calls, API traces, or screenshot resemblance, and a low score under these proxies cannot say whether the model saw poorly, planned poorly, or was failed by its harness. We study scientific overview figure reconstruction, an agent task in which a source image must become an editable PowerPoint slide that preserves text, topology, layout, and native document structure. We introduce ReFigBench, a benchmark and evaluation framework built on 1,000 real overview figures retrieved from arXiv papers with full provenance. Coding agents from four model families reconstruct every figure under two workflows, direct code generation and a specialized PPTX workflow, and the strongest model runs inside two commercial harnesses, yielding ten configurations. Evaluation combines deterministic artifact checks, repeated automated scoring by judges from two model families, and blinded human comparisons. Perception remains a bottleneck that iterative rendering only partly repays. Whether workflow effort converts into quality depends on the model together with its harness, since the same model gains from the specialized workflow inside one harness and loses inside the other, and the harness shifts scores even under an identical direct prompt. The specialized workflow erases native connectors in every configuration, human judges still prefer its renderings in most matchups, and even the strongest agent falls short of the rubric ceiling. These results expose the tension between fidelity and editability as the central challenge for practical multimodal document agents.
Comments31 pages, 7 figures, including appendices