Multimodal large language models and physics visual tasks: comparative analysis of performance and costs
专题命中 其他VLM :multimodal large language model(title,abstract)
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 其他VLM :multimodal large language model(title,abstract)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI
机构 * Department of Computer Science and Engineering, IIT Bombay(印度理工学院班加罗尔计算机科学与工程系) ; Center of Machine Intelligence and Data Science (C-MInDS), IIT Bombay(印度理工学院班加罗尔人工智能与数据科学中心)
专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI
机构 * School of Computer Science ; Engineering University of New South Wales Sydney, Australia ; Stanford University California, USA ; Computer Science University College London London, UK ; School of Eng. Maths. \& Tech University of Bristol Bristol, UK ; Inst. Logic Language \& Computation University of Amsterdam Amsterdam, NL
专题命中 其他VLM :vision-language model(abstract);分类 cs.AI
Comments Accepted to: 2025 IEEE International Conference on Quantum Artificial Intelligence (QAI), Naples, Italy, Nov 2-5, 2025. This is the authors' accepted manuscript (AAM). An IEEE copyright notice appears on page 1. The final published version will appear in IEEE Xplore; DOI to be added when available
机构 * MERaLiON Team(MERaLiON团队) ; Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI
机构 * National Taiwan Normal University(台湾国立台湾师范大学)
专题命中 其他VLM :multimodal large language model(abstract)
Comments submitted to the ISCA SLaTE-2025 Workshop