CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
Hongyong Han, Wei Wang, Gaowei Zhang, Mingjie Li, Yi Wang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Technology Innovation Center for South China Sea Remote Sensing, Surveying and Mapping Collaborative Application, Ministry of Natural Resources(自然资源部南海遥感测绘协同应用技术创新中心)
;
South China Sea Development Research Institute, Ministry of Natural Resources(自然资源部南海发展研究 institute)
;
Inspur Computer Technology Co., Ltd(Inspur 计算机技术有限公司)
;
Shandong Key Laboratory of Advanced Computing(山东先进计算重点实验室)
CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
Qiangguo Jin, Xianyao Zheng, Hui Cui, Changming Sun, Yuqi Fang, Cong Cong, Ran Su, Leyi Wei, Ping Xuan, Junbo Wang
机构
*
School of Software, Northwestern Polytechnical University, Shaanxi, China(西北工业大学软件学院)
;
Yangtze River Delta Research Institute of Northwestern Polytechnical University, Taicang, China(西北工业大学长江三角研究 institute)
;
Department of Computer Science and Information Technology, La Trobe University, Melbourne, Australia(拉筹伯大学计算机科学与信息技术系)
;
CSIRO Data61, Sydney, Australia(CSIRO Data61)
;
School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院)
;
Australian Institute of Health Innovation (AIHI), Macquarie University, Australia(麦考瑞大学健康创新研究所)
;
School of Computer Software, College of Intelligence and Computing, Tianjin University, Tianjin, China(天津大学计算机软件学院)
;
Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University, Macao Special Administrative Region of China(澳门理工学院人工智能驱动药物发现中心)
;
Department of Computer Science, School of Engineering, Shantou University, Guangdong, China(汕头大学计算机科学系)
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Peking University(北京大学)
;
Tencent(腾讯)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Tsinghua University(清华大学)
;
Nanyang Technological University(南洋理工大学)
;
Westlake University(西湖大学)
专题命中
视觉问答
:multimodal large language model(title,abstract);分类 cs.CV
BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities
Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Sara Pieri, Saeed Yahya Alseiari, Shanavas Cholakkal, Khaled Aldahmani, Fahad Khan, Rao Anwer, Salman Khan, Timothy Baldwin, Hisham Cholakkal
机构
*
Mohamed Bin Zayed University of Artificial Intelligence(Mohamed Bin Zayed大学人工智能学院)
;
Linköping University(林肯大学)
;
Shaikh Tahnoon bin Mohammed Medical City(Shaikh Tahnoon bin Mohammed医疗城)
;
Tawam Hospital(Tawam医院)
;
Sheikh Shakhbout Medical City(Sheikh Shakhbout医疗城)
;
Govt Medical College Kozhikode(科钦政府医学院)
机构
*
Nanjing University(南京大学)
;
Xiaohongshu Inc.(小红书公司)
;
East China Normal University(华东师范大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Oxford(牛津大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Trishanu Das, Abhilash Nandy, Khush Bajaj, Deepiha S
机构
*
Tredence Inc.(特伦德公司)
;
Indian Institute of Technology Kharagpur(印度理工学院克拉格浦尔分校)
;
Inria Paris-Rocquencourt(巴黎-罗克琴库特研究所)
;
Rajiv Gandhi University(拉吉夫·甘地大学)
;
Tsinghua University(清华大学)
;
Palmer Research Laboratories(帕勒姆研究实验室)
Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models
Junjie Chen, Xuyang Liu, Subin Huang, Linfeng Zhang, Hang Yu
机构
*
Anhui Polytechnic University (AHPU)(安徽工程大学)
;
Shanghai University (SHU)(上海大学)
;
Shanghai Jiao Tong University (SJTU)(上海交通大学)
;
Sichuan University (SCU)(四川大学)
机构
*
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济发展实验室)
;
SKL-IOTSC, CIS, University of Macau(澳门科学馆物联网与智慧城市研究中心)
;
Shandong University(山东大学)
;
Auckland University of Technology(技术大学(奥克兰))
;
University of Electronic Science and Technology of China(电子科技大学)
OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery
Xiucheng Liang, Jinheng Xie, Tianhong Zhao, Rudi Stouffs, Filip Biljecki
机构
*
Department of Architecture, National University of Singapore(建筑系,新加坡国立大学)
;
Department of Electrical and Computer Engineering, National University of Singapore(电气与计算机工程系,新加坡国立大学)
;
School of Artificial Intelligence, Shenzhen Technology University(人工智能学院,深圳科技大学)
;
Department of Real Estate, National University of Singapore(房地产系,新加坡国立大学)
专题命中
视觉定位与Grounding
:vision-language model(abstract);VLM(abstract);multimodal large language model(abstract);分类 cs.CV
Journal refISPRS Journal of Photogrammetry and Remote Sensing 230: 918-942, 2025
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
University of Washington(华盛顿大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
CommentsAccepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)