机构
*
School of Computer and Communication Sciences, EPFL(瑞士联邦理工学院计算机与通信科学学院)
;
Microsoft Research Asia(微软亚洲研究院)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
机构
*
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学)
;
Department of Orthopedics, Qilu Hospital, Shandong University(骨科部,齐鲁医院,山东大学)
;
Key Laboratory of Qingdao in Medicine and Engineering(医学与工程青岛重点实验室)
Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models
Junjie Chen, Xuyang Liu, Subin Huang, Linfeng Zhang, Hang Yu
机构
*
Anhui Polytechnic University (AHPU)(安徽工程大学)
;
Shanghai University (SHU)(上海大学)
;
Shanghai Jiao Tong University (SJTU)(上海交通大学)
;
Sichuan University (SCU)(四川大学)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar
机构
*
ServiceNow
;
York University(约克大学)
;
Mila – Quebec AI Institute(魁北克人工智能研究院)
;
École de Technologie Supérieure(魁北克高等技术学院)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
University of Waterloo(滑铁卢大学)
;
CIFAR AI Chair(CIFAR人工智能 chair)
;
Polytechnique Montréal(蒙特利尔理工学院)
;
University of British Columbia(不列颠哥伦比亚大学)
$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Trishanu Das, Abhilash Nandy, Khush Bajaj, Deepiha S
机构
*
Tredence Inc.(特伦德公司)
;
Indian Institute of Technology Kharagpur(印度理工学院克拉格浦尔分校)
;
Inria Paris-Rocquencourt(巴黎-罗克琴库特研究所)
;
Rajiv Gandhi University(拉吉夫·甘地大学)
;
Tsinghua University(清华大学)
;
Palmer Research Laboratories(帕勒姆研究实验室)
机构
*
School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院)
;
Miner School of Computer & Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校Miner计算机与信息科学学院)
OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery
Xiucheng Liang, Jinheng Xie, Tianhong Zhao, Rudi Stouffs, Filip Biljecki
机构
*
Department of Architecture, National University of Singapore(建筑系,新加坡国立大学)
;
Department of Electrical and Computer Engineering, National University of Singapore(电气与计算机工程系,新加坡国立大学)
;
School of Artificial Intelligence, Shenzhen Technology University(人工智能学院,深圳科技大学)
;
Department of Real Estate, National University of Singapore(房地产系,新加坡国立大学)
专题命中
图文多模态
:multimodal(abstract);分类 cs.CV
Journal refISPRS Journal of Photogrammetry and Remote Sensing 230: 918-942, 2025