ACE:用于多幻灯片演示自动化的自修正智能体画布编辑器
ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation
- Seoul National University(首尔大学)
- Miridih(米里迪)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM智能体编辑演示文档的布局破坏与有效输出被误判问题,提出带CARE路由器和自修正循环的ACE,其性能优于对比方法,获盲评者偏好。
AI中文摘要:
商业设计平台越来越多地通过大语言模型(LLM)智能体编辑文档,但可靠部署存在两个实际问题: legacy文档格式仅暴露平面化、绝对定位的元素,因此智能体必须重新计算坐标,且常破坏布局;设计没有唯一的 ground truth,因此与参考对比的指标会惩罚有效但不同的输出。我们提出ACE,这是一种基于分层场景图(hierarchical scene-graph)的智能体画布编辑器,具有面向演示的动作空间(98种工具),搭配CARE(一种内容感知路由器,仅向智能体提供每张幻灯片的相关片段,平均减少约89%的输入token),以及由无 ground truth 的指令遵循(IF)评判器驱动的自修正循环,该评判器的自然语言批评会作为下一轮指令反馈给智能体。在固定主干下,单轮的场景图编辑器已与同主干的内部迭代智能体HTML流水线表现相当;加入自修正后,ACE在指令遵循任务上显著优于该流水线(94个任务的完整基准上IF得分4.23 vs 3.81,配对p=0.010,经外部评判器复现),速度为其1.75倍,成本低约44%。VQ均值无统计差异,但26名盲评者整体更偏好ACE(58.7%的决定性胜率),且81%的情况下偏好自修正后的输出;该排名在三类评判器中保持一致,外部评判器保留了三分之二的自修正增益,限制了循环性。66%的案例在一轮后停止,严格峰值回滚消除了所有观察到的退化。
英文摘要:
Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only \emph{flat}, absolutely positioned elements, so agents must recompute coordinates and routinely break layouts; and design has no unique ground truth, so diff-against-reference metrics penalize valid-but-different outputs. We present \textbf{ACE}, an agentic canvas editor over a \emph{hierarchical scene-graph} with a presentation-specialized action space (98 tools), paired with \textbf{CARE}, a content-aware router that feeds the agent only the relevant slice of each deck (avg.\ $\sim$89\% input-token reduction), and a \emph{self-correction} loop driven by a \emph{ground-truth-free} instruction-following (IF) judge whose natural-language critique is fed back as the next-turn instruction. With a fixed backbone, a scene-graph editor in a \emph{single turn} already matches a same-backbone \emph{agentic} HTML pipeline that iterates internally; adding self-correction lifts ACE significantly above it on instruction following (IF 4.23 vs.\ 3.81 on the full 94-task benchmark, paired $p{=}.010$, replicated by an out-of-loop judge) at 1.75$\times$ the speed and $\sim$44\% lower cost. VQ means are statistically indistinguishable, but 26 blind raters prefer ACE overall (58.7\% decisive win-rate) and prefer the self-corrected output 81\% of the time; the ranking is invariant across three judge families, and out-of-loop judges retain two-thirds of the self-correction gain, bounding circularity. 66\% of cases halt after one pass, and a strict-peak rollback removes every observed regression.