arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29718cs.CLcs.AI

PPTBench:编码智能体能否通过结构化、可编辑的幻灯片重建视觉世界?

PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides

Xiaoqiu Wang, Yizhe Chi, Wenyi Li, Deyao Hong, Zhihan Shan, Mingju Gao, Kaisen Yang, Youjie Zheng, Calvin Xiao, Qinhuai Na

首次发表
浏览论文内容

中文总结 AI 辅助

PPTBench通过500个基于arXiv论文流程图的可编辑幻灯片重建任务,评估编码智能体的视觉编码能力,发现当前最佳配置Kimi K3得分仅67.80,智能体在语义和视觉细节上仍有不足。

中文摘要 AI 辅助

编码智能体正开始涉足视觉世界。它们现在能够构建网页、图形用户界面、游戏、3D场景、图表和文档。此类视觉编码的成功需要连接两个空间:推断视觉结构并以编程方式表达。幻灯片是知识工作的核心媒介,广泛用于以人们可直接检查和编辑的形式交流想法和协作。因此,它们为视觉编码提供了理想的测试平台,因为这要求智能体恢复视觉结构并将其实现为可编辑对象。然而,现有基准要么依赖主观的开放式评估,要么产生不可编辑的代码输出,要么仅关注局部编辑而非端到端的视觉重建。我们引入了PPTBench,它通过可编辑的幻灯片重建来基准测试视觉编码。它包含500个任务,每个任务基于来自真实arXiv论文的科学流程图,要求智能体将其重建为包含原生可编辑对象的单个PPTX页面。一个四阶段的Agentic Judge评估工件有效性、语义正确性、渲染质量和细粒度视觉质量。在涵盖模型家族、努力程度和框架的31种配置中,最佳配置Kimi K3仅达到67.80,而中位数得分为19.47。我们发现智能体能够可靠地生成有效的PPTX文件,但在语义和视觉正确性方面仍存在困难,尤其是文本细节。更多的推理主要帮助智能体通过严格的关卡,而更强的验证与更高质量更一致地相关。PPTBench推进了编码智能体通过结构化、可编辑代码理解和重建视觉世界的愿景。

英文摘要

Coding agents are increasingly moving beyond text-based software tasks to reconstruct visual targets through code. This capability, visual coding, requires agents to translate their understanding of visual targets into executable code. Slides provide a natural testbed for this capability, combining rich visual structure with objects that can be programmatically created and edited. To measure this capability, we introduce PPTBench, a benchmark for reconstructing scientific flow diagrams as editable PowerPoint slides. PPTBench covers 500 scientific flow-diagram tasks across 10 presentation domains, drawn from real research papers, and is evaluated with a four-stage agentic judge covering artifact validity, process and connector fidelity, rendering quality, and visual fidelity. Evaluation of ten models across 36 model--harness--effort configurations shows a substantial gap in reliable visual coding: the best configuration achieves 77.34, while the median across configurations is 24.38. Fine-grained analysis shows that agents can generally produce valid slide files, but struggle to produce high-quality reconstructions that faithfully recover the semantics and visual structure of the target. Further analysis shows that increasing reasoning effort primarily improves hard-gate passage rather than mean detail quality on each configuration's passing tasks, while configurations with more inspection tend to achieve higher overall scores. PPTBench establishes a measurable testbed for studying visual coding and advancing agents toward more reliable visual creation.

发表机构

  • Navers Lab, Einsia.AI(Navers实验室,Einsia.AI)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑