发表机构
Universidad Industrial de Santander; King Abdullah University of Science and Technology (KAUST)(桑坦德工业大学; 阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对教育视频合成中空间正确性和视觉清晰度难题,提出符号几何智能体SGA,通过拦截代码、提取场景图及针对性优化提升质量,引入MVQS,实验显示其在多配置下有效提升了视觉质量分数。
AI 中文摘要
近期工作利用大语言模型(LLMs)为教学动画生成可执行代码。但确保空间正确性和视觉清晰度仍具挑战,因现有框架忽视几何遮挡。我们提出符号几何智能体(SGA),用于以代码为中心的动画管道,拦截LLM生成的代码,执行部分提取符号场景图,检测冲突时应用针对性优化。还引入Manim视觉质量分数(MVQS)。实验表明SGA在多个配置中提升了MVQS。
英文摘要
Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim. However, ensuring spatial correctness and visual legibility remains challenging, as existing frameworks emphasize pedagogical content while overlooking geometric occlusions. We propose the Symbolic Geometric Agent (SGA), a plug-and-play module for code-centric animation pipelines that intercepts LLM-generated code, performs partial execution to extract symbolic scene graphs, and applies targeted refinement when spatial conflicts are detected. We further introduce the Manim Visual Quality Score (MVQS), a deterministic rendering-free proxy for spatial integrity. Experiments on the MMMC-Code benchmark across four LLM backbones and two agentic pipelines show that SGA achieves a peak MVQS of 73.11 (Code2Video + GPT-5.1), corresponding to a 16.1% relative improvement over the raw baseline, and improves MVQS in 7 of 8 backbone x pipeline configurations. Additionally, we conduct a human evaluation showing that these improvements translate into human preference, with raters preferring SGA over the raw baseline in 84.4% of comparisons and over a VLM-based critic in 65.0%.