arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21455cs.CV

CompAdapt:面向物理一致文本到视频生成的自适应复合运动建模

CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation

Haoran Qin, Renlong Wu, Tianyu Huang, Yukang Ding, Hui Li, Wangmeng Zuo

首次发表
浏览论文内容

中文总结 AI 辅助

CompAdapt提出物理一致文本到视频生成框架,扩展至复合运动,通过结构化语义和动力学感知先验匹配实现自适应,提升物理一致性与视觉质量。

中文摘要 AI 辅助

虽然基于扩散的文本到视频(T2V)模型在生成逼真且时间连贯的视频方面展现了令人印象深刻的能力,但它们往往无法遵循基本的物理动力学。尽管最近的物理约束方法引入了显式动力学先验以提高物理合理性,但它们仍局限于简单的单一类型运动,依赖手动指定的参数,并且难以泛化到未见过的物理定律。在这项工作中,我们提出了CompAdapt,一个物理一致的T2V框架,用于在复杂的现实世界场景中进行自适应生成。它将神经动力学建模从单一类型运动扩展到涵盖复合物理行为,包括耦合运动、多阶段转换和多物体碰撞。此外,CompAdapt将自然语言提示转换为结构化的物理语义,实现了运动类型、时间关系和初始物理参数的端到端指定。为了泛化到新颖的物理环境,CompAdapt引入了动力学感知先验匹配,实现了无需重新训练核心动力学模块的一次性自适应。此外,一个物理感知的潜在特征融合模块提高了在快速和复杂运动下的视觉保真度。在物理聚焦的T2V基准上的实验表明,CompAdapt在物理一致性上优于一般的T2V模型和物理约束基线,同时保持了高视觉质量和对未见动力学的适应性。项目页面可在以下https URL获取。

英文摘要

While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, they remain limited to simple single-type motions, depend on manually specified parameters, and struggle to generalize to unseen physical laws. In this work, we propose CompAdapt, a physics-consistent T2V framework for adaptable generation across complex real-world scenarios. It extends neural dynamics modeling beyond single-type motions to encompass composite physical behaviors, including coupled motions, multi-stage transitions, and multi-object collisions. Furthermore, CompAdapt translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial physical parameters. To generalize to novel physical environments, CompAdapt introduces dynamics-aware prior matching, achieving one-shot adaptation without retraining the core dynamics module. In addition, a physics-aware latent feature fusion module improves visual fidelity under fast and complex motion. Experiments on physics-focused T2V benchmarks demonstrate that CompAdapt improves physical consistency over both general T2V models and physics-constrained baselines, while preserving high visual quality and adaptability to unseen dynamics. The project page is available at https://makapic.github.io/CompAdapt/ .

发表机构

  • Harbin Institute of Technology(哈尔滨工业大学)
  • Taobao, Alibaba Group(阿里巴巴集团淘宝)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑