How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?
机构 * Department of Mathematics and Computer Science(数学与计算机科学系)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Department of Mathematics and Computer Science(数学与计算机科学系)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
机构 * School of Computation, Information and Technology, and the School of Medicine and Health(计算信息学院及医学健康学院) ; Technical University of Munich(慕尼黑技术大学) ; Department of Computing(计算系) ; Institut de Neurosciences de la Timone, UMR 7289, CNRS, Aix-Marseille Université(神经科学研究所,UMR 7289,CNRS,艾克斯-马赛大学) ; Aix-Marseille Univ, APHM, Service de Neuroradiologie Diagnostique et Interventionnelle, Hôpital de la Timone(艾克斯-马赛大学,APHM,诊断与介入神经放射科,泰米翁医院) ; Aix-Marseille Univ, APHM, service de neurologie pédiatrique, Hôpital de la Timone(艾克斯-马赛大学,APHM,儿童神经科,泰米翁医院)
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV
Comments Work currently under revision for IEEE TMI
机构 * Department of Information and Communication Engineering(信息与通信工程系)
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV
机构 * Tsinghua University(清华大学) ; University of Science and Technology of China(中国科学技术大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Comments Accepted to CVPR
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Honda Research Institute USA(本田研究院美国)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Comments Paper is accepted by IJCV
机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) ; School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) ; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; Meituan(美团)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by CVPR2025
机构 * The Hong Kong University of Science & Technology (Guangzhou)(香港科技大学(广州)) ; The Hong Kong University of Science & Technology(香港科技大学) ; Zhejiang University(浙江大学) ; Lappeenranta-Lahti University of Technology(拉佩兰塔-拉赫蒂技术大学)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted by CVPR 2025
机构 * University of Waterloo(滑铁卢大学) ; Votee AI ; Vector Institute(向量研究所)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments accepted to ACL 2025 main, camera ready
机构 * Intelligent Computing and Machine Learning Lab, School of ASEE, Beihang University(北京航空航天大学自动化学院智能计算与机器学习实验室) ; Xiaohongshu(小红书) ; School of Sino-French Engineer, Beihang University(北京航空航天大学中法工程师学院) ; College of Engineering and Computer Science, VinUniversity(Vin大学工程与计算机科学学院)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
机构 * College of Application and Technology, Shenzhen University, China(应用技术学院,深圳大学,中国) ; College of Big Data and Internet, Shenzhen Technology University, China(大数据与互联网学院,深圳科技大学,中国)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
Comments NLDB 2025
专题命中 视频多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.MM
Comments Accepted by ICMR 2025
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Comments CVPR 2025
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV
Comments 11 pages, 5 figures
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments 5 pages, 2 figures
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments CVPR 2025, Project Page: https://videoautoarena.github.io/
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
Comments NAACL 2025 Main
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Comments Accepted to AAAI2025
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by ICASSP 2025
专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments 8 pages, 5 figures, submitted to ACM Multimedia Asia 2024
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments 11pages, 5figs
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted by TCSVT 2024
专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV
Comments Extension of CVPR'23 paper for journal submission
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV
Comments Accepted for publication at the ECCV 2024 workshop on Neuromorphic Vision: Advantages and Applications of Event Cameras (NEVI)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL