CinemaTraj:用LLM智能体为3D场景合成原子相机轨迹
CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents
浏览论文内容
中文总结 AI 辅助
CinemaTraj是将LLM智能体与3D场景图结合的框架,可从自然语言生成带旁白的3D场景相机轨迹,在真实环境中优于现有方法。
中文摘要 AI 辅助
从自然语言描述自动生成3D场景中具有电影表现力的相机轨迹是一项极具实用价值的挑战性任务,应用范围涵盖房地产广告到虚拟游览创建。现有方法要么依赖2D图像先验缺乏真正的3D空间感知,要么将轨迹生成视为脱离电影语义的几何路径规划问题。我们提出CinemaTraj,该框架将相机轨迹规划重新定义为语言 grounding 的空间推理问题。给定一组RGB-D图像和用户提示,CinemaTraj为LLM智能体配备结构化3D场景图:智能体将提示分解为一系列原子电影运动(推拉、环绕、升降、平移、俯仰、变焦、弧形运动)。每个运动通过一种新颖的参数化轨迹表示实现,该表示兼具电影表现力与避障优化能力。场景图作为结构化空间先验,将智能体的推理建立在环境的准确几何与语义知识之上。CinemaTraj还生成与相机运动同步的旁白和字幕,产生带旁白的电影视频输出。我们在真实世界ScanNet++环境上评估CinemaTraj,结果表明它能生成符合提示要求、无碰撞且具有高电影质量的轨迹,在提示对齐、轨迹质量和安全性指标上均优于现有方法。
英文摘要
Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high practical value, with applications ranging from real-estate advertising to virtual tour creation. Existing methods either lack true 3D spatial awareness by relying on 2D image priors, or treat trajectory generation as a geometric path planning problem divorced from cinematographic semantics. We present CinemaTraj, a framework that reframes camera trajectory planning as a language-grounded spatial reasoning problem. Given a set of RGB-D images and a user prompt, CinemaTraj equips an LLM agent with a structured 3D scene graph: the agent decomposes the prompt into a sequence of atomic cinematographic movements (dolly, orbit, crane, pan, tilt, zoom, arc). Each movement is instantiated via a novel parametric trajectory representation that is both cinematographically expressive and optimizable for collision avoidance. The scene graph acts as a structured spatial prior, grounding the agent's reasoning in accurate geometric and semantic knowledge of the environment. CinemaTraj further generates synchronized voiceover and subtitles aligned with camera motion, producing narrated cinematic video outputs. We evaluate CinemaTraj on real-world ScanNet++ environments, and show that it produces prompt-faithful, collision-free trajectories with high cinematographic quality, outperforming existing approaches on prompt alignment, trajectory quality, and safety metrics.
发表机构
- Technical University of Munich(慕尼黑工业大学)
- Huawei Dresden Research Center(华为德累斯顿研究中心)
机构由 AI 辅助整理,请以论文原文为准。