AI 中文总结
VectorHarness是一个多智能体框架,旨在将科学图形重建为原生可编辑表示,通过类型合适的对象恢复和联合评估基准,提高编辑成功率与关系保持,同时减少光栅回退并保持高视觉保真度。
AI 中文摘要
将科学图形转换为可编辑表示仍然是图像到代码生成中的一个挑战性问题,因为其元素异构且布局复杂。最近的多智能体重建系统推进了这项工作,但通常遵循复制-粘贴范式:重建图像与原始图像高度相似,而复杂区域实际上仍不可编辑。我们反而提出了一个不同的目标,即光栅到作者重建,旨在恢复支持原生、定制编辑而非仅仅视觉复制的作者表示。为此,我们提出了VectorHarness,一个用于光栅到作者重建的多智能体框架,使用类型合适的原生表示来恢复异构组件。文本、公式、形状、连接器、图标、图表和表格被重建为原生可编辑对象,而本质上基于图像的区域则保留为光栅内容。为了系统评估重建质量,我们引入了VectorHarness-Bench,它联合评估渲染保真度、光栅回退覆盖率、可执行对象编辑和保持关系的编辑。实验表明,VectorHarness提高了可执行编辑成功率和关系保持,减少了可避免的光栅回退,并在异构图形中保持了高视觉保真度。
英文摘要
Converting scientific graphics into editable representations remains a challenging problem for image-to-code generation because of their heterogeneous elements and complex layouts. Recent multi-agent reconstruction systems have advanced this line of work, but often follow a copy-paste paradigm: the reconstructed image closely resembles the original, while complex regions remain effectively uneditable. We instead formulate a different objective, raster-to-authoring reconstruction, which aims to recover an authoring representation that supports native, customized editing rather than mere visual replication. To this end, we present VectorHarness, a multi-agent framework for raster-to-authoring reconstruction that recovers heterogeneous components using type-appropriate native representations. Text, formulas, shapes, connectors, icons, charts, and tables are reconstructed as natively editable objects, while intrinsically image-based regions remain raster content. To systematically evaluate reconstruction quality, we introduce VectorHarness-Bench, which jointly assesses rendering fidelity, raster fallback coverage, executable object edits, and relation-preserving edits. Experiments show that VectorHarness improves executable edit success and relation preservation, reduces avoidable raster fallback, and maintains high visual fidelity across heterogeneous graphics.