PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
PhyVLLM:具有运动-外观解耦的物理引导视频语言模型
机构 * Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) ; Department of Electronic Engineering, Tsinghua University(电子工程系,清华大学)
专题命中 动作与事件理解 :video language model(title);video understanding(abstract);video-language(abstract);分类 cs.CV
AI总结 PhyVLLM通过引入物理运动建模和运动-外观解耦,提升视频语言模型在物理推理和视频理解任务中的性能。