Chunze Yang, Qidong Liu, Wenjie Zhao, Yue Tang, Jiusong Ge, Di Zhang, Jiashuai Liu, Lei Wu, Junbo Lu, Ni Zhang, Xian Wu, Zeyu Gao, Chen Li
机构
*
School of Comp. Science & Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
;
Tencent Jarvis Lab(腾讯Jarvis实验室)
;
University of Cambridge(剑桥大学)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models
重新审视遥感中的变化视觉问答问题:结构化与原生多模态Qwen模型
Yakoub Bazi, Mohamad M. Al Rahhal, Mansour Zuair, Faroun Mohamed
机构
*
Computer Engineering Department, College of Computer and Information Sciences, King Saud University(计算机工程系,计算机与信息科学学院,沙特国王大学)
;
Applied Computer Science Department, College of Applied Computer Science, King Saud University(应用计算机科学系,应用计算机科学学院,沙特国王大学)
MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage
MedObvious:通过临床分诊暴露视觉语言模型中的医学莫拉维奇悖论
Ufaq Khan, Umair Nawaz, L D M S S Teja, Numaan Saeed, Muhammad Bilal, Yutong Xie, Mohammad Yaqub, Muhammad Haris Khan
机构
*
Mohamed bin Zayed University of Artificial Intelligence, UAE(马尔代夫布扎伊德人工智能大学,阿联酋)
;
National Institute of Technology, Silchar(西尔CHAR国家理工学院)
;
Birmingham City University, UK(伯明翰城市大学,英国)
专题命中
视觉问答
:vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI
Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
Patho-R1: 基于多模态强化学习的病理专家推理器
Wenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu
机构
*
Department of Pathology, West China Hospital, Sichuan University(四川大学华西医院病理科部门)
;
Institute of Clinical Pathology, West China Hospital, Sichuan University(四川大学华西医院临床病理科研究所)
;
University of Toronto(多伦多大学)
;
Business School, Sichuan University(四川大学商学院)
;
Department of Pathology, Shengjing Hospital of China Medical University(中国医科大学盛京医院病理科部门)
专题命中
视觉问答
:vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
M4-RAG:大规模多语言多文化多模态检索增强生成
David Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee, Genta Indra Winata
机构
*
Stanford University(斯坦福大学)
;
MBZUAI
;
Indian Institute of Science(印度科学研究院)
;
Ontario Tech University(安大略技术大学)
;
University of Toronto(多伦多大学)
;
Capital One
AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions
AQuA:迈向具有模糊性视觉问题的策略性响应生成
Jihyoung Jang, Hyounghun Kim
机构
*
Graduate School of Artificial Intelligence, POSTECH(人工智能研究生院,POSTECH)
;
Department of Computer Science and Engineering, POSTECH(计算机科学与工程系,POSTECH)
Remote Sensing Retrieval-Augmented Generation: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model
遥感检索增强生成:通过多模态数据集和检索增强生成模型连接遥感图像与综合知识
Congcong Wen, Yiting Lin, Xiaokang Qu, Nan Li, Yong Liao, Xiang Li, Hui Lin
机构
*
School of Cyber Science and Technology, University of Science and Technology of China(信息科学技术学院,中国科学技术大学)
;
China Academy of Electronics and Information Technology(电子信息技术研究院)
Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective Prediction
利用数据说不:基于记忆的插拔式选择预测
Aditya Sarkar, Yi Li, Jiacheng Cheng, Shlok Mishra, Nuno Vasconcelos
机构
*
University of Maryland, College Park(马里兰大学)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
Qualcomm AI(高通人工智能)
;
Yale University(耶鲁大学)
;
Meta AI