arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24208cs.CV

CoaG:网格上的圆柱体:用于视频生成的粗略3D布局控制

CoaG: Cylinders on a Grid for Coarse 3D Layout Control in Video Generation

Zhangsihao Yang, Mengyi Shan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出CoaG方法,仅用地面网格和人物圆柱体即可控制视频中人物位置与摄像机运动,通过自动生成训练数据并微调LoRA,实现遵循布局和路径的逼真视频生成。

中文摘要 AI 辅助

我们探究一个人需要绘制多少几何图形才能控制生成视频中人物的站立位置和摄像机的移动。我们的答案是:一个地面平面和每个人一个圆柱体。用户在地面上绘制一个网格,在每个应站立的位置放置一个圆柱体,在81帧中移动圆柱体和摄像机,模型便渲染出一个逼真的视频,其中人物占据圆柱体的位置,随圆柱体的移动而移动,并从所绘制的摄像机视角被看到。外观来自文本提示和背景参考图像;布局和运动来自几何图形。由于没有数据集将此类信号与视频配对,我们自行构建配对:一个自动引擎从组合种子生成2000条字幕,为每条字幕使用文本到视频模型生成一个片段,并通过人物跟踪、背景修复、智能体地面掩码循环、前馈多视图重建和平面拟合将每个片段提升回其几何图形,全程无真实镜头和手动标签。在Wan2.2-Fun-Control上训练的LoRA,基于1935个这样的元组,在保留片段上遵循绘制的布局和摄像机路径:生成的人物与圆柱体的数量、顺序、位置和高度匹配,文本改变他们的身份,参考图像改变他们的位置,并且推进、环绕、平移和升降路径被遵循,只有拉远路径较弱。

英文摘要

We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81 frames, and the model renders a photoreal video in which the people occupy the cylinders' positions, move as the cylinders move, and are seen from the drawn camera. Appearance comes from a text prompt and a background reference image; layout and motion come from the geometry. Because no dataset pairs such a signal with video, we build the pairs ourselves: an automatic engine writes 2000 captions from a combinatorial seed, generates a clip for each with a text-to-video model, and lifts every clip back to its geometry with person tracking, background inpainting, an agentic ground-mask loop, feed-forward multi-view reconstruction and a plane fit, with no real footage and no manual labels. A LoRA on Wan2.2-Fun-Control trained on 1935 such tuples follows drawn layouts and camera paths on hold-out clips: the generated people match the cylinders' count, order, position and height, the text changes who they are, the reference image changes where they are, and dolly-in, orbit, pan and crane paths are followed, dolly-out only weakly.

发表机构

  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑