SeaSlides:面向智能体幻灯片生成的语义抽象层
SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation
AI总结:
研究针对智能体幻灯片生成的现有系统适配性差问题,提出SeaSlides框架,通过语义抽象层结合HTML与Typst后端,在多模型评估中提升了丰富内容生成质量,验证了语义抽象原则的有效性。
AI中文摘要:
智能体演示文稿生成必须保留源内容、维持连贯视觉设计、渲染专用对象并生成可用成品。现有系统仅满足部分要求:模板可保持规整性但限制适应性,而自由格式HTML或SVG赋予模型灵活性的代价是需处理底层渲染决策。这种不匹配导致长技术幻灯片脆弱,尤其当幻灯片包含公式、代码或数据图表时。我们提出SeaSlides,一个围绕语义抽象层构建的智能体幻灯片生成框架。模型无需编写坐标、内联样式或原始SVG几何,而是通过可复用组件和能力模块编写结构化幻灯片内容,同时模板负责布局、样式和渲染。我们在HTML和Typst中分别实现该原理:SeaSlides-HTML使用模板定义的DOM组件,SeaSlides-Typst使用模板函数和包支持的模块。能力模块将公式、代码和图表路由至专用渲染器,三个反馈阶段在导出前定位构建错误、项目约束违规和视觉缺陷。两个系统保留后端特定语法和契约,同时共享相同创作边界。为评估,我们将128任务的UltraPresent验证设置与新的32任务基准SeaSlidesBench-Rich结合,该基准侧重数学、代码、伪代码、表格、图表和图形。在四个生成模型上,两个SeaSlides后端生成的源文件比依赖SVG的生成更具可读性和内容导向性。SeaSlides后端在四个模型中的三个上达到最高的丰富内容宏平均,同时保持有竞争力的整体定性性能。这些结果支持语义抽象作为跨演示后端的实用创作原则。
英文摘要:
Agentic presentation generation must preserve source content, maintain coherent visual design, render specialized objects, and produce usable artifacts. Existing systems meet only part of this requirement: templates preserve regularity but restrict adaptation, whereas free-form HTML or SVG gives models flexibility at the cost of low-level rendering decisions. This mismatch makes long technical decks brittle, especially when slides contain formulas, code, or data graphics. We present SeaSlides, an agentic slide-generation framework built around a semantic abstraction layer. Rather than authoring coordinates, inline styles, or raw SVG geometry, the model writes structured slide content through reusable components and capability modules, while templates own layout, style, and rendering. We instantiate this principle separately in HTML and Typst: SeaSlides-HTML uses template-defined DOM components, whereas SeaSlides-Typst uses template functions and package-backed modules. Capability modules route equations, code, and charts to dedicated renderers, and three feedback stages localize build errors, project-constraint violations, and visual defects before export. The two systems retain backend-specific syntax and contracts while sharing the same authoring boundary. For evaluation, we combine the 128-task UltraPresent validation setting with SeaSlidesBench-Rich, a new 32-task benchmark stressing mathematics, code, pseudocode, tables, charts, and diagrams. Across four generation models, both SeaSlides backends produce more readable, content-oriented source than SVG-heavy generation. A SeaSlides backend attains the highest rich-content macro-average under three of the four models while maintaining competitive overall qualitative performance. These results support semantic abstraction as a practical authoring principle across presentation backends.