CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
机构 * Sharif University of Technology(谢里夫理工大学)
专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI
Comments Accepted to TMLR 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Sharif University of Technology(谢里夫理工大学)
专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI
Comments Accepted to TMLR 2025
机构 * Department of Biomedical Engineering, University of Basel, Allschwil, Switzerland(巴塞尔大学生物医学工程系) ; Clarunis – University Digestive Health Care Center Basel, Basel, Switzerland(Clarunis – 巴塞尔大学消化健康医疗中心)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 13 pages, 3 figures; accepted at ML-CDS @ MICCAI 2025, Daejeon, Republic of Korea
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
专题命中 视频多模态 :multimodal(title);MLLM(abstract);分类 cs.CV
机构 * School of Electronic Information, Wuhan University, Wuhan, China(武汉大学电子信息学院) ; MiLM Plus, Xiaomi Inc., Wuhan, China(小米公司MiLM Plus团队)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
专题命中 视频多模态 :multi-modal(title,abstract)
机构 * Allen Wang Gavin Tao
专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments 14pages, 9 figures, Journal paper
机构 * School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院) ; DAMO Academy, Alibaba group(阿里巴巴集团大模型研究院) ; Hupan Lab(虎扑实验室) ; National Key Laboratory of General Artificial Intelligence, Peking University(北京人工智能 general artificial intelligence 国家重点实验室) ; Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI
Comments ICCV 2025. Code: https://github.com/Gen-Verse/Paper2Video
专题命中 视频多模态 :multimodal(abstract);分类 cs.AI、cs.MM
Comments Accepted to the Proceedings of the 2025 International Conference on Artificial Intelligence and Virtual Reality (AIVR 2025). \c{opyright} 2025 Springer. This is the author-accepted manuscript. Rui Xi and Xianghan Wang contributed equally to this work. The final version will be available via SpringerLink
机构 * University of Twente(特文特大学) ; University of Bath(巴斯大学) ; Shanghai Jiao Tong University(上海交通大学) ; PhiGent Robotics(PhiGent机器人)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments Accepted by ICRA 2025
机构 * Keye Team, Kuaishou Group(快手集团Keye团队)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments Github page: https://github.com/Kwai-Keye/Keye
机构 * Zhejiang University and Alibaba Cloud(浙江大学和阿里云) ; Alibaba Cloud(阿里云) ; Zhejiang University(浙江大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments Accepted by T-ASE and CoRL25 GenPriors Workshop