SLIDEFORGE:一款用于将幻灯片作为结构化工件进行可控编辑的大语言模型智能体
SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts
浏览论文内容
中文总结 AI 辅助
SLIDEFORGE是一款基于Deck State Graph的LLM智能体,可实现保留布局样式等的可控幻灯片编辑,在多维度评估中优于现有基线,提供可控幻灯片编辑的新方案。
中文摘要 AI 辅助
当前AI智能体能够出色地描述幻灯片,但AI辅助幻灯片编辑不仅需要理解,还要求输出保留布局、样式、组件结构和原生可编辑性。现有用于AI辅助幻灯片编辑的智能体基于截图或弱文档表示运行,常将连贯视觉单元碎片化、栅格化可编辑内容或破坏布局。相比之下,针对可控幻灯片编辑,我们提出智能体框架SLIDEFORGE,其构建了Deck State Graph(幻灯片状态图),这是一种可执行的幻灯片状态,将视觉分解、原生pptx对象结构与感知组织关联起来。通过恢复可被人类参考的组件同时保留细粒度可编辑结构,SLIDEFORGE支持通过幻灯片原生操作实现主题保留重构,并进行渲染状态验证。我们进一步提出可控幻灯片转换的评估范式,该范式联合衡量组件恢复、保留、重样式一致性、视觉质量和原生可编辑性。实验表明,SLIDEFORGE在这些维度上的表现优于直接提示、基于截图的智能体和通用代码智能体基线。代码可在该https URL获取。
英文摘要
Current AI agents compellingly describe slides. However, AI-assisted slide editing requires more than understanding: the output must retain layout, style, component structure, and native editability. Towards, AI-assisted slide editing, existing agents operate on screenshots or weak document representations and often fragment coherent visual units, rasterize editable content, or break layout. In contrast, for controllable slide editing, we introduce an agentic framework, SLIDEFORGE, which builds a Deck State Graph, an executable slide state that links visual decomposition, native pptx object structure, and perceptual organization. By recovering human-referable components while retaining fine-grained editable structure, SLIDEFORGE supports theme-preserving reconstruction through slide-native operations and rendered-state verification. We further introduce an evaluation paradigm for controllable slide transformation that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability. Experiments show that SLIDEFORGE outperforms direct prompting, screenshot-based agents, and generic code-agent baselines across these dimensions. Code is available at https://github.com/UIUC-MONET/SLIDEFORGE.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Meta
机构由 AI 辅助整理,请以论文原文为准。