Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
机构 * SETLabs Resarch GmbH(SETL实验室)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 11 pages, 3 figures
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * SETLabs Resarch GmbH(SETL实验室)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 11 pages, 3 figures
机构 * Institute of Neuroinformatics, University of Zurich(神经信息学研究所,苏黎世大学) ; ETH Zurich(苏黎世联邦理工学院) ; Digital Society Initiative, University of Zurich(数字社会倡议,苏黎世大学) ; Department of Geography(地理系)
专题命中 视频多模态 :multimodal(title,abstract)
Comments 8 pages
机构 * School of Computer Science and Technology, Anhui University(计算机科学与技术学院,安徽大学)
专题命中 视频多模态 :multi-modal(abstract,comments);cross-modal(abstract);分类 cs.CV、cs.AI
Comments The First Work that Exploits Multi-modal Knowledge Graph for Pedestrian Attribute Recognition
机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所模式识别实验室) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Kling Team, Kuaishou Technology(快手科技 Kling 团队) ; Peking University(北京大学) ; Nanjing University(南京大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) ; School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments *Corresponding authors: Min Zhu (min.zhu@scu.edu.cn) and Junlong Cheng (jlcheng@scu.edu.cn)
机构 * HKUST (GZ)(香港科技大学(广州)) ; HKUST(香港科技大学) ; HIT(哈尔滨工业大学)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV