机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学医学智能与XR研究所)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
ProMSA: 渐进式多模态搜索智能体用于基于知识的视觉问答
ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen, Yunyao Yu, Chuanrui Zhang, Zirui Liao, Jun Yang, Zhenyu Yang, Haonan Lu, Haoqian Wang
机构
*
OPPO AI Center, OPPO Inc. China(OPPO AI中心,OPPO公司)
;
The Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
AI 是否能像艺术史家一样看?解析视觉语言模型如何识别艺术风格
Marvin Limpijankit, Milad Alshomary, Yassin Oulad Daoud, Amith Ananthram, Tim Trombley, Emily L. Spratt, Anna Filonenko, Hannah Pivo, Elias Stengel-Eskin, Mohit Bansal, Noam M. Elcott, Kathleen McKeown
机构
*
Columbia University, Department of Computer Science(哥伦比亚大学计算机科学系)
;
Columbia University, Department of Art History & Archaeology(哥伦比亚大学艺术史与考古系)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
UNC Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中
视觉问答
:vision language model(title);visual question answering(abstract);分类 cs.CV、cs.AI
V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views
V2X-QA:面向自动驾驶多视角的多模态大语言模型综合数据集与基准测试
Junwei You, Pei Li, Zhuoyu Jiang, Weizhe Tang, Zilin Huang, Rui Gan, Jiaxi Liu, Yan Zhao, Sikai Chen, Bin Ran
机构
*
Department of Civil and Environmental Engineering, University of Wisconsin–Madison(威斯康星大学麦迪逊分校土木与环境工程系)
;
Department of Civil and Architectural Engineering and Construction Management, University of Wyoming(怀俄明大学土木与建筑工程及施工管理系)
;
School of Transportation, Southeast University(东南大学交通学院)
专题命中
视觉问答
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
TPCL:面向鲁棒视觉问答的任务渐进课程学习
Ahmed Akl, Abdelwahed Khamis, Zhe Wang, Ali Cheraghian, Sara Khalifa, Kewen Wang
机构
*
School of Information and Communication Technology(信息与通信技术学院)
;
Data61, CSIRO Australia(Data61,澳大利亚联邦科学与工业研究组织)
;
School of Information Systems(信息系统学院)
机构
*
1 Institute of Automation, Chinese Academy of Sciences
;
2 School of Artificial Intelligence, University of Chinese Academy of Sciences
;
3 Centre for Artificial Intelligence
;
Robotics, Hong Kong Institute of Science \& Innovation, CAS
;
4 Independent Researcher
Saliency Guided Longitudinal Medical Visual Question Answering
基于显著性的纵向医学视觉问答
Jialin Wu, Xiaofeng Liu
机构
*
Dept. of Computer Science and Engineering University of California, San Diego(计算机科学与工程系,加州大学圣地亚哥分校)
;
Dept. of Radiology and Biomedical Imaging Yale University(放射学与生物医学成像系,耶鲁大学)
CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
Qiangguo Jin, Xianyao Zheng, Hui Cui, Changming Sun, Yuqi Fang, Cong Cong, Ran Su, Leyi Wei, Ping Xuan, Junbo Wang
机构
*
School of Software, Northwestern Polytechnical University, Shaanxi, China(西北工业大学软件学院)
;
Yangtze River Delta Research Institute of Northwestern Polytechnical University, Taicang, China(西北工业大学长江三角研究 institute)
;
Department of Computer Science and Information Technology, La Trobe University, Melbourne, Australia(拉筹伯大学计算机科学与信息技术系)
;
CSIRO Data61, Sydney, Australia(CSIRO Data61)
;
School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院)
;
Australian Institute of Health Innovation (AIHI), Macquarie University, Australia(麦考瑞大学健康创新研究所)
;
School of Computer Software, College of Intelligence and Computing, Tianjin University, Tianjin, China(天津大学计算机软件学院)
;
Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University, Macao Special Administrative Region of China(澳门理工学院人工智能驱动药物发现中心)
;
Department of Computer Science, School of Engineering, Shantou University, Guangdong, China(汕头大学计算机科学系)