LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) ; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及交叉应用关键实验室(东南大学),中华人民共和国教育部,中国) ; School of Software, Northwestern Polytechnical University(西北工业大学软件学院) ; State Key Laboratory for Novel Software Technology and School of Intelligence Science and Technology, Nanjing University(新型软件技术国家重点实验室和南京大学智能科学与技术学院) ; School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
Comments Accepted by TPAMI, Code is available at: https://github.com/wengwanjiang/FoundSkelModel
机构 * Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) ; University of Adelaide(阿德莱德大学)
专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted to ICCV 2025
机构 * Lawrence Berkeley National Laboratory(伯克利国家实验室) ; University of California, Irvine(加州大学尔湾分校) ; University of California, Berkeley(加州大学伯克利分校) ; Covalent Metrology(协力计量)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments This paper has been accepted for presentation at the 59th International Conference on Parallel Processing (ICPP 2025), DRAI workshop
专题命中 视频多模态 :MLLM(abstract);分类 cs.CV