发表机构
Institute of Mathematics of the Romanian Academy; National University of Science and Technology Politehnica Bucharest; Büchi Labortechnik AG(罗马尼亚科学院数学研究所; 布加勒斯特理工大学; 布赫实验室技术公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GEST引擎从自然语言文本生成多角色视频,核心是明确的世界模型,以GEST表示世界状态。通过程序或智能系统生成GEST,经四阶段管道执行,能输出多种视频相关数据,且保证相关特性,其输出可作视频理解多方面工具。
AI 中文摘要
我们展示了GEST引擎,这是一个从自然语言文本到完全注释的多角色视频的完整系统。其核心是一个明确的世界模型:该引擎不是将状态编码为学习到的潜在表示,而是维护世界的完整、可检查表示(存在哪些角色、它们在哪里、在做什么、持有哪些物体以及事件在时空上如何关联),表示为时空事件形式图(GEST),并在通过开源多人脚本框架驱动的商业游戏引擎的开放世界中确定性地实现。GEST可以通过程序生成,也可以由智能文本到GEST系统生成,其中大语言模型主管通过程序状态后端验证的工具调用规划一个故事,因此每个生成的规范在构建时都是可执行的。然后GEST进入一个四阶段执行管道:图解析和验证、实体和动作基础、时间编排(通过Floyd-Warshall传递闭包解决Allen风格的约束)以及执行和捕获。在一次模拟过程中,该引擎以零边际注释成本发出帧对齐的RGB视频、每像素深度密集图、实例分割、每个角色的骨骼姿势、每帧成对空间关系图、二维边界框、事件到帧的时间映射以及自然语言描述。我们还描述了一个游戏内世界编辑器运行时能力提取、文本生成管道以及一个跨并行虚拟机大规模渲染语料库的生产系统。由于每个帧都可追溯到语义规范,该引擎通过构建保证了对象永久性多角色协调和时间一致性,使其输出作为视频理解的训练数据、评估基准和诊断工具很有价值。
英文摘要
We present the GEST-Engine, a complete system that goes from natural-language text to fully-annotated multi-actor video. At its core is an explicit world model: rather than encoding state as a learned latent, the engine maintains a complete, inspectable representation of the world (which actors exist, where they are, what they are doing, which objects they hold, and how events relate in time and space), expressed as a formal Graph of Events in Space and Time (GEST) and realized deterministically inside the open world of a commercial game engine driven through an open-source multiplayer scripting framework. GESTs are produced either procedurally or by an agentic text-to-GEST system in which an LLM Director plans a story through tool calls validated by a programmatic state backend, so every generated specification is executable by construction. A GEST then enters a four-stage execution pipeline: graph parsing and validation, entity and action grounding, temporal orchestration (Allen-style constraints resolved by Floyd-Warshall transitive closure), and execution and capture. In a single simulation pass the engine emits frame-aligned RGB video, dense per-pixel depth, instance segmentation, per-actor skeletal pose, per-frame pairwise spatial-relation graphs, 2D bounding boxes, event-to-frame temporal mappings, and natural-language descriptions, all at zero marginal annotation cost. We further describe an in-game world editor, runtime capability extraction, a text-generation pipeline, and a production system that renders corpora at scale across parallel virtual machines. Because every frame traces back to a semantic specification, the engine guarantees object permanence, multi-actor coordination, and temporal consistency by construction, making its output valuable as training data, evaluation benchmarks, and diagnostic tools for video understanding.