PPTBench:编码智能体能否通过结构化、可编辑的幻灯片重建视觉世界?
PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides
浏览论文内容
中文总结 AI 辅助
PPTBench通过500个基于arXiv论文流程图的可编辑幻灯片重建任务,评估编码智能体的视觉编码能力,发现当前最佳配置Kimi K3得分仅67.80,智能体在语义和视觉细节上仍有不足。
中文摘要 AI 辅助
编码智能体正开始涉足视觉世界。它们现在能够构建网页、图形用户界面、游戏、3D场景、图表和文档。此类视觉编码的成功需要连接两个空间:推断视觉结构并以编程方式表达。幻灯片是知识工作的核心媒介,广泛用于以人们可直接检查和编辑的形式交流想法和协作。因此,它们为视觉编码提供了理想的测试平台,因为这要求智能体恢复视觉结构并将其实现为可编辑对象。然而,现有基准要么依赖主观的开放式评估,要么产生不可编辑的代码输出,要么仅关注局部编辑而非端到端的视觉重建。我们引入了PPTBench,它通过可编辑的幻灯片重建来基准测试视觉编码。它包含500个任务,每个任务基于来自真实arXiv论文的科学流程图,要求智能体将其重建为包含原生可编辑对象的单个PPTX页面。一个四阶段的Agentic Judge评估工件有效性、语义正确性、渲染质量和细粒度视觉质量。在涵盖模型家族、努力程度和框架的31种配置中,最佳配置Kimi K3仅达到67.80,而中位数得分为19.47。我们发现智能体能够可靠地生成有效的PPTX文件,但在语义和视觉正确性方面仍存在困难,尤其是文本细节。更多的推理主要帮助智能体通过严格的关卡,而更强的验证与更高质量更一致地相关。PPTBench推进了编码智能体通过结构化、可编辑代码理解和重建视觉世界的愿景。
英文摘要
Coding agents are increasingly moving beyond text-based software tasks to reconstruct visual targets through code. This capability, visual coding, requires agents to translate their understanding of visual targets into executable code. Slides provide a natural testbed for this capability, combining rich visual structure with objects that can be programmatically created and edited. To measure this capability, we introduce PPTBench, a benchmark for reconstructing scientific flow diagrams as editable PowerPoint slides. PPTBench covers 500 scientific flow-diagram tasks across 10 presentation domains, drawn from real research papers, and is evaluated with a four-stage agentic judge covering artifact validity, process and connector fidelity, rendering quality, and visual fidelity. Evaluation of ten models across 36 model--harness--effort configurations shows a substantial gap in reliable visual coding: the best configuration achieves 77.34, while the median across configurations is 24.38. Fine-grained analysis shows that agents can generally produce valid slide files, but struggle to produce high-quality reconstructions that faithfully recover the semantics and visual structure of the target. Further analysis shows that increasing reasoning effort primarily improves hard-gate passage rather than mean detail quality on each configuration's passing tasks, while configurations with more inspection tend to achieve higher overall scores. PPTBench establishes a measurable testbed for studying visual coding and advancing agents toward more reliable visual creation.
发表机构
- Navers Lab, Einsia.AI(Navers实验室,Einsia.AI)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。