arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25321cs.AI

基于模拟数据集和双流光流监督的物理基础流体视频生成

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

Ruijie Su, Yuanzhi Liang, Xiaohua Xie, Jianhuang Lai

首次发表
浏览论文内容

中文总结 AI 辅助

研究流体视频生成中模型违反物理原理的问题,构建含模拟与真实视频的数据集,引入双流图像到视频架构,通过光流监督增强模型,提升视频物理常识和质量得分,使模型内化连贯运动先验。

中文摘要 AI 辅助

视频扩散模型生成的视觉内容令人信服,但在处理流体主题时常常违反基本物理原理,如液柱在空气中破裂、倒水时容器水位不上升等。我们认为这是因为大规模视频文本语料库几乎没有明确的运动监督,模型学习模仿流体外观而非动力学。为此,我们有两项贡献。一是构建了一个物理模拟流体数据集,包含多种视频及测试集。二是引入双流图像到视频架构,基于预训练扩散变压器视频生成器,通过光流解码器分支增强标准RGB解码器,并将其损失融合到RGB流中。在两个模型规模和两个测试集上,我们的方法在视频物理常识和视频质量得分上有显著提升,优于竞争对手,且在盲测中受人类评分者青睐。直接光流读出评估显示分布内端点误差低至0.54像素,证实模型内化了连贯运动先验而非仅改善表面外观。

英文摘要

Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity. We attribute this gap to the fact that large-scale video-text corpora contain almost no explicit motion supervision, so models learn to imitate fluid appearance rather than dynamics. We address this with two contributions. First, we build a physics-simulation fluid dataset combining 1,638 MPM-simulated pouring/sloshing videos with 2,320 keyword-filtered real pouring videos mined from stock footage, plus two held-out test sets: a 1,515-video real-video benchmark and an 18-prompt text-to-first-frame generalization benchmark. Second, we introduce a dual-stream image-to-video architecture built on a pretrained diffusion-transformer video generator. It augments the standard RGB decoder with a lightweight Optical-Flow Decoder branch trained with explicit end-point-error and smoothness losses, fused into the RGB stream via zero-initialized convolutions so the pretrained backbone starts undisturbed. Only the two decoders are updated; the encoder, temporal transformer, and text encoder remain frozen. Across two model scales (1.3B and 14B) and two test sets, our method improves VideoPhy-2 Physical-Commonsense and Video-Quality scores over the frozen backbone by up to 8.75 and 4.65 points, outperforms a leading open competitor, and is preferred by human raters in a blind study. A direct optical-flow read-out evaluation further shows an end-point error as low as 0.54 pixels in-distribution, confirming the model has internalized a coherent motion prior rather than merely improving surface appearance.

发表机构

  • Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑