Enhancing Multi-Image Question Answering via Submodular Subset Selection
机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔学院)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔学院)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) ; Key Laboratory of Social Computing and Cognitive Intelligence (Dalian University of Technology), Ministry of Education, China(社会科学计算与认知智能重点实验室(大连理工大学)) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV
Comments 9 pages, 5 figures
机构 * East China Normal University(东华师范大学) ; Midea Group, AI Lab(美的集团人工智能实验室) ; Syracuse University(雪城大学) ; Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) ; Shanghai University(上海大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments add more citations
机构 * Amazon Web Services(亚马逊网络服务)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments N/A
机构 * Department of Electronic and Communication Engineering, North China Electric Power University(电子与通信工程系,华北电力大学) ; Hebei Key Laboratory of Power Internet of Things Technology, North China Electric Power University(河北省电力物联网技术重点实验室,华北电力大学) ; Computer Science and Software Engineering Department, Monmouth University(计算机科学与软件工程系,蒙特莫恩大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
机构 * HKUST(香港科技大学) ; University of Waterloo(滑铁卢大学) ; Vector Institute(向量研究所)
专题命中 图文多模态 :multimodal(abstract);分类 cs.AI
Comments Preprint
机构 * School of Computer Engineering & Science, Shanghai University(上海大学计算机工程与科学学院) ; Tencent YouTu Lab(腾讯优图实验室)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV
Comments Accepted to CVPR 2025
机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学) ; Aalto University(艾尔沃斯大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
Comments Accepted at ACL 2024-BIONLP Workshop. Code: https://github.com/mbzuai-oryx/XrayGPT
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments 9 pages,4 figures, 4 tables,Innovation in Medicine and Healthcare Proceedings of 13th KES-InMed 2025
机构 * Department of Computer Science and Engineering, University at Buffalo(计算机科学与工程系,布法罗大学) ; Department of Surgery, University at Buffalo(外科系,布法罗大学) ; Department of Computer Science and Engineering, IIT, Jodhpur(计算机科学与工程系,印度理工学院,乔杜尔) ; Department of Computer Science and Engineering, University of Notre Dame(计算机科学与工程系,圣母大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
机构 * Department of Electrical Engineering and Computer Sciences, University of California-Berkeley, Berkeley, CA, USA(电气工程与计算机科学系,加州大学伯克利分校)
专题命中 图文多模态 :multimodal(abstract);分类 cs.AI
Comments 28 pages, 10 figures, 19 tables
机构 * Purdue University(普渡大学) ; Clarkstown High School South(Clarkstown 高中南分校) ; University at Albany, State University of New York(阿尔巴尼大学,纽约州立大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV
专题命中 图文多模态 :multimodal(abstract);分类 cs.AI
Comments PhD thesis, 123 pages
机构 * AI, HKUST(GZ)(香港科技大学(广州)人工智能学院) ; ZJUT(浙江工业大学) ; CSE, HKUST(香港科技大学计算机科学与工程系)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV
机构 * Intel Labs(英特尔实验室) ; National Research Council Canada(加拿大国家研究理事会)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted to NAACL 2025 main track (oral)
机构 * Meta FAIR ; Fudan University(复旦大学) ; Meta Reality Labs
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Updated refs, fixed typos, and added new COCO SotA: 66.0 val mAP! Code, models, and data at https://github.com/facebookresearch/perception_models
机构 * School of Computer Science and Technology, Guangdong University of Technology(广东技术大学计算机科学与技术学院) ; School of Computer Science and Intelligence Education, Lingnan Normal University(岭南师范学院计算机科学与智能教育学院)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments 10 pages
机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(先进交叉学科研究院,中国科学院大学) ; Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV
机构 * Cornell University(康奈尔大学) ; University of Peloponnese(希腊皮洛斯大学) ; Kansas State University(堪萨斯州立大学) ; Universidad de las Fuerzas Armadas(武装力量大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
机构 * Nanyang Technological University(南洋理工大学) ; University of Oxford(牛津大学) ; Tel Aviv University(特拉维夫大学) ; MILA(蒙特利尔人工智能研究院) ; ERA-Krueger AI Safety Lab(ERA-Krueger人工智能安全实验室)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Published at ICLR 2025
机构 * The University of Hong Kong(香港大学) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Project page: https://zcmax.github.io/projects/LLaVA-3D/
机构 * Department of Medicine, University of Florida(医学系,佛罗里达大学) ; Department of Electrical and Computer Engineering, University of Florida(电气与计算机工程系,佛罗里达大学) ; Department of Biomedical Engineering, University of Florida(生物医学工程系,佛罗里达大学) ; Intelligent Clinical Care Center, University of Florida(智能临床护理中心,佛罗里达大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学) ; Alibaba Group(阿里巴巴集团) ; UIUC(伊利诺伊大学香槟分校)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CL
机构 * Information Technology University(信息科技大学) ; University of Ontario Institute of Technology(Ontario Institute of Technology 大学) ; Mohamed Bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Added code link in the abstract
机构 * School of Computer Science, Nanjing University of Information Science & Technology(信息科学技术南京大学计算机科学学院) ; Jiangsu Collaborative Innovation Center of Atmospheric Environment and Equipment Technology (CICAEET)(江苏省大气环境与装备技术协同创新中心) ; School of Computer Science and Informatics, Cardiff University(计算机科学与信息学院,卡迪夫大学) ; Faculty of Information Technology and Communication Sciences, Tampere University(信息科技与通信科学学院,坦佩雷大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV
专题命中 图文多模态 :multimodal(abstract);分类 cs.CL
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV