MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
机构 * School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) ; Zhongguancun Academy(中关村学院) ; Southeast Academy of Information Technology, Beijing Institute of Technology(信息技术东南学院,北京理工大学) ; Tsinghua University(清华大学)
专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV
Comments Under Review