OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments 9 pages, 6figures
AI 大模型
视频理解、视频生成、视频语言模型和时序视觉推理。
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments 9 pages, 6figures
机构 * Rice University(里士大学)
专题命中 视频生成 :text-to-video(title,abstract);分类 cs.CV
Comments Added new video experiments and more image experiments to validate the method
机构 * Xi’an Key Laboratory of Big Data and Intelligent Vision, Xidian University, Xi’an 710071, China(西安大数据与智能视觉重点实验室,西安电子科技大学,西安710071,中国) ; Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University, Xi’an 710071, China(协同智能系统重点实验室,教育部,西安电子科技大学,西安710071,中国) ; School of Computer Science and Technology, Xidian University, Xi’an 710071, China(计算机科学与技术学院,西安电子科技大学,西安710071,中国)
专题命中 视频生成 :video generation(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学) ; University of California, Los Angeles(加州大学洛杉矶分校) ; Palo Alto Networks(帕洛阿尔托网络公司) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 视频扩散模型 :text-to-video(title,abstract);video diffusion(title);分类 cs.CV
Comments To appear in the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP)
机构 * ByteDance Intelligent Creation(字节跳动智能创作)
专题命中 视频扩散模型 :video generation(title);分类 cs.CV
机构 * The University of British Columbia(不列颠哥伦比亚大学) ; University of Toronto(多伦多大学) ; Vancouver General Hospital, Vancouver, BC, Canada(温哥华总医院)
专题命中 视频扩散模型 :video diffusion(title)
Comments Data Curation and Augmentation in Medical Imaging CVPR 2024