DrawAI:用于生成可编辑光栅图像的智能体基准与工作流
DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable
浏览论文内容
中文总结 AI 辅助
DrawAI提出图像到可编辑重建任务,构建含DrawAI-Bench基准与DrawAI-Flow工作流的智能体方案,经实验验证DrawAI-Flow可提升可编辑结构,不同模型-harness配置的重建质量与成本差异显著。
中文摘要 AI 辅助
近期,图像生成模型和多模态智能体可针对日益复杂的视觉通信任务生成高质量视觉内容,但其光栅输出仍难以直接使用,因为有意义的内容与关系被压缩为像素,导致用户无法检查、修改、重新排列或重用单个组件。我们提出图像到可编辑重建任务,旨在从光栅图像中恢复结构化、可直接操作的产物,同时保留其视觉与语义内容,核心挑战是同时满足保真度与可编辑性,而两者在实际应用中常存在权衡。为研究该任务,我们引入DrawAI,包含智能体基准DrawAI-Bench与重建工作流DrawAI-Flow。DrawAI-Bench涵盖科学图表、演示幻灯片、海报和示意图,结合真实图像与AI生成图像以反映实际视觉创作场景,通过含39项指标的混合协议评估保真度与可编辑性:确定性基于规则的指标测量具有直接对应关系的属性,特定资产的视觉语言评分标准捕捉精确匹配易产生误导的语义与感知质量。此外,我们提出DrawAI-Flow,这是一个两阶段智能体工作流,其中解析智能体将提取的元素证据转化为显式重建计划,重建智能体通过迭代的代码-渲染-验证-修正循环将计划实现为可执行图形代码。在DrawAI-Bench上,我们针对五个智能体 harness 系统评估了十三个模型,以研究模型能力、harness 选择与工作流设计的影响。结果显示,不同模型-harness 配置的重建质量与成本差异显著,而DrawAI-Flow可持续提升可编辑结构。
英文摘要
Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their raster outputs remain difficult to use directly because meaningful content and relationships are flattened into pixels, preventing users from inspecting, modifying, rearranging, or reusing individual components. We formulate image-to-editable reconstruction, which recovers a structured, directly manipulable artifact from a raster image while preserving its visual and semantic content. The central challenge is to jointly satisfy Fidelity and Editability, which often trade off in practice. To study this task, we introduce DrawAI, comprising an agentic benchmark, DrawAI-Bench, and a reconstruction workflow, DrawAI-Flow. DrawAI-Bench spans scientific figures, presentation slides, posters, and diagrams, combining real and AI-generated images to reflect practical visual-creation scenarios. It evaluates Fidelity and Editability through a hybrid protocol of 39 criteria: deterministic rule-based metrics measure properties with direct correspondences, while asset-specific vision-language rubrics capture semantic and perceptual qualities for which exact matching is misleading. Besides, we propose DrawAI-Flow, a two-stage agentic workflow in which a Parser Agent turns extracted elements evidence into an explicit reconstruction plan, and a Reconstruction Agent realizes the plan as executable graphics code through an iterative code-render-validate-revise loop. On DrawAI-Bench, we systematically evaluate thirteen models across five agent harnesses to study the effects of model capability, harness choice, and workflow design. The results show that reconstruction quality and costs vary substantially across model-harness configurations, while DrawAI-Flow consistently improves editable structure.
发表机构
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Tsinghua University(清华大学)
- Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。