发表机构
KAIST AI(韩国科学技术院人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出EditPPT多智能体框架,通过PowerPoint COM接口执行局部操作并结合双模态验证,在DeckEdit-Bench基准上实现高执行率等指标,解决长文档幻灯片编辑的级联错误问题。
AI 中文摘要
自动化幻灯片编辑需同时满足修改准确性、内容保真度及对文档长度的鲁棒性。现有基于大语言模型(LLM)的系统在处理实际演示文件时常失效,因其依赖理想化中间表示或开放式代码生成,易在长文档中产生级联错误。本文提出EditPPT,一种将幻灯片编辑重新表述为受约束工具选择问题的多智能体框架。通过PowerPoint原生COM接口执行局部形状级操作,EditPPT缩小LLM动作空间,同时保留用户创建文档的应用解析结构。通过跨模态分离验证,本文的双模态验证器能更鲁棒地评估指令保真度与视觉质量。本文还推出DeckEdit-Bench基准,包含28个人工创建的文档、582张幻灯片及183个编辑提示,覆盖短、中、长文档层级。实验显示,EditPPT整体实现99.5%的执行率、88.7%的幻灯片定位F1值、82.5%的指令遵循度及91.5%的对象保留率,且在长文档上保持优异性能。代码与基准可在指定URL获取。
英文摘要
Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Existing LLM-based systems often fail on real-world presentation files because they rely on idealized intermediate representations or open-ended code generation, which are prone to cascading errors in long decks. We introduce EditPPT, a multi-agent framework that reformulates slide editing as a constrained tool-selection problem. By executing localized shape-level operations through the native PowerPoint COM interface, EditPPT narrows the LLM action space while preserving the application-resolved structure of user-authored decks. By separating validation across modalities, our dual-modal validation provides more robust assessment of both instruction fidelity and visual quality. We also present DeckEdit-Bench, a benchmark with 28 human-authored decks, 582 slides, and 183 editing prompts across short, medium, and long deck tiers. Experiments show that EditPPT achieves a 99.5% execution rate, 88.7% slide-targeting F1, 82.5% instruction following, and 91.5% object preservation overall, while maintaining strong performance on long decks. Our code and benchmark are available at https://anonymous.4open.science/r/EditPPT-0E27/
Comments30 pages, 7 figures, 17 tables, EMNLP 2026 submitted, under review