Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
机构 * The Chinese University of Hong Kong(香港中文大学)
专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments Accepted by NeurIPS 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * The Chinese University of Hong Kong(香港中文大学)
专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments Accepted by NeurIPS 2025
专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI
机构 * Mission San Jose High School(Mission San Jose 高中) ; Missouri University of Science and Technology(密苏里科学与技术大学) ; Michigan State University(密歇根州立大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
机构 * Sanya Science and Education Innovation Park, Wuhan University of Technology(武汉理工大学三亚科学教育创新园) ; Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology(湖北省交通运输物联网重点实验室,计算机科学与人工智能学院,武汉理工大学) ; State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
机构 * CUHK MMLab(香港中文大学多模态实验室) ; CUHK (SZ)(香港中文大学(深圳)) ; Tsinghua University(清华大学) ; UCAS(中国科学院大学) ; CUHK HCCL(香港中文大学高性能计算实验室)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments NeurIPS 2025, Project page: https://github.com/tulerfeng/Video-R1