arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PACE:精确AI电影化表达——面向脚本预可视化与几何一致性的类型化规范

PACE: Precise AI Cinematic Expression

Bing Duan, Qiang Guo, Linpu Li, Zhijian Mao, Min Zhu, Zhirui Ren, Yiwei Yan, Xi Chu, Xiaoding Li

arXiv 2609.19853首次发表:更新:

发表机构

Studio π(Studio π)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PACE提出一种类型化规范,将剧本转化为精确的摄像机与场景布局,通过编译器生成提示词和3D场景,并逐字段测量几何一致性,实验表明其能显著提升动作绘制准确率。

AI 中文摘要

在剧本与电影之间,存在一个以空间为首要的规划问题:谁站在哪里,以及摄像机从其所处位置能看到什么。当以自由文本形式向图像扩散模型请求一个镜头时,该模型会依据其自身默认设置来解决这一规划问题。我们提出PACE(精确AI电影化表达),一种用于该规划的类型化表示:剧本依据、所需的角色、道具和地点、每个主体所站的位置,以及摄像机的动作。一个值在其所属层级(剧本、场景、镜头或分镜)被写入一次,并向下继承。一个编译器将结果转化为发送给扩散模型的提示词,以及一个以米为单位构建的3D场景;一个摄像机求解器放置摄像机,使得所声明的取景即为所构建的取景。当声明的值成为几何时,PACE逐字段地测量编译后的摄像机与舞台化渲染结果偏离声明的程度,而非要求模型进行判断。在包含11个场景的Automatic Drive剧本上,每个单主体的舞台化分镜将其主体放置在距其声明位置不超过帧宽1.2%的范围内;当有两个或三个主体时,单一的摄像机姿态无法满足所有位置,此时会报告残差而非将其吸收。在204个外部导演-故事板镜头中,从导演文字得到的交付头部高度是舞台化目标的1.906倍,从编译提示词得到的是1.733倍,而使用灰盒控制则为0.955倍;最能保持取景的条件所绘制的动作最少。在30个镜头上声明姿态,将所绘制的动作从58.9%提升至74.4%,且不改变取景。转场、拟合运动以及人工审查生成的画面仍是待办事项。代码:此HTTPS链接

英文摘要

Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the plan: the screenplay evidence, the characters, props and locations it needs, where each subject stands, and what the camera does. A value is written once at the level it belongs to (script, scene, shot or panel) and inherited below it. A compiler turns the result into both the prompt sent to the diffusion model and a 3D scene built in metres, and a camera solver places the camera so that the declared framing is the framing built. Where a declared value becomes geometry, PACE measures, field by field, how far the compiled camera and the staged render sit from the declaration, rather than asking a model to judge. On the 11-scene Automatic Drive screenplay, every staged single-subject panel places its subject within 1.2% of frame width of its declared position; with two or three subjects one camera pose cannot satisfy every position, and the residual is reported rather than absorbed. On 204 external director-storyboard shots, delivered head height is 1.906 times the staged target from the director's words, 1.733 from the compiled prompt, and 0.955 with the greybox control; the condition that holds framing best draws the described action least. Declaring the pose on 30 shots raises the action drawn from 58.9% to 74.4% without moving the framing. Transitions, fitted motion and human review of the generated panels remain open. Code: https://github.com/StudioPiLabs/pace-core

Commentsv2: the supplementary material referenced throughout v1 was never uploaded; it is removed and its 69 references resolved, two of its results moved into the main text and one dropped. 36 pages, 8 figures, 4 tables. Code: https://github.com/StudioPiLabs/pace-core

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑