arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21580cs.CVcs.AI

GraphVid:交互式图可控视频生成

GraphVid: Interactive Graph-Controllable Video Generation

Vedant Shah, Onkar Susladkar, Tushar Prakash, Kiet Nguyen, Tianjiao Yu, Adheesh Juvekar, Muntasir Wahed, Ismini Lourentzou

首次发表
浏览论文内容

中文总结 AI 辅助

研究如何实现可控视频生成,提出基于图条件的GraphVid模型及相关数据集,通过结构化交互图实现灵活精确控制,相比其他方法在减少训练数据和参数的情况下,提升了视频生成的可控性与质量。

中文摘要 AI 辅助

可控视频生成具有挑战性,因为难以用文本提示或主要约束像素移动的运动控制输入来指定精确的多对象交互。基于轨迹的控制存在问题。为此引入GraphVid,一种基于图条件的图像到视频生成模型,通过结构化交互图实现交互式控制。还构建了GraphVid - Bench数据集。尽管训练数据和可训练参数更少,但GraphVid具有很强的可控性和视频质量,与Motion - I2V相比有显著提升。结果凸显了结构化语义接口在可控视频生成中的潜力。

英文摘要

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce $\textbf{GraphVid}$, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate $\textbf{GraphVid-Bench}$, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.

↑