TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
André G. Viveiros, Patrick Fernandes, Saul Santos, Sonal Sannigrahi, Emmanouil Zaranis, Nuno M. Guerreiro, Amin Farajian, Pierre Colombo, Graham Neubig, André F. T. Martins
机构
*
Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术高级学院)
;
Instituto de Telecomunicações(电信研究所)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Sword Health(Sword健康)
;
TransPerfect
;
MICS, CentraleSupélec, Université Paris-Saclay(MICS,中央圣艾尔布里大学,巴黎萨克雷大学)
;
ELLIS Unit Lisbon(里斯本ELLIS单位)
机构
*
Northeastern University(东北大学)
;
Microsoft Research(微软研究院)
;
University of Southern California(南加州大学)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models
Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari, Mingkun Xu
机构
*
Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院)
;
Agency for Science, Technology and Research(科技研究局)
;
School of Information Technology(信息技术学院)
Modest-Align: Data-Efficient Alignment for Vision-Language Models
Jiaxiang Liu, Yuan Wang, Jiawei Du, Joey Tianyi Zhou, Mingkun Xu, Zuozhu Liu
机构
*
Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院)
;
ZJU-Angelalign R&D Center for Intelligence Healthcare(浙大天使align智能医疗研发中心)
;
Centre for Frontier AI Research (CFAR)(前沿人工智能研究中心)
;
Agency for Science, Technology and Research (A*STAR)(科技研究局)
;
Institute of High Performance Computing (IHPC)(高性能计算研究所)
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
Lu Zhang, Jiazuo Yu, Haomiao Xiong, Ping Hu, Yunzhi Zhuge, Huchuan Lu, You He
机构
*
Dalian University of Technology(大连理工大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Tsinghua Shenzhen International Graduate School(清华大学深圳国际graduate school)
EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking
Yukuan Zhang, Jiarui Zhao, Shangqing Nie, Jin Kuang, Shengsheng Wang
机构
*
College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)
;
Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(吉林大学教育部长春符号计算与知识工程重点实验室)
;
Yangtze University(扬子大学)
机构
*
Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
;
Hong Kong Polytechnic University(香港理工大学)
;
Zhuhai Campus of Sun Yat-sen University(中山大学珠海校区)
;
University of Adelaide(阿德莱德大学)
;
Peking University(北京大学)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis
Feng Yuan, Yifan Gao, Yuehua Ye, Haoyue Li, Xin Gao
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Suzhou Institute of Biomedical Engineering and Technology(苏州生物医学工程与技术研究所)
;
Chinese Academy of Sciences(中国科学院)
;
The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院)
机构
*
School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院)
;
College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院)
;
School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室)
;
School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室)
;
Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)
Calibration-Aware Prompt Learning for Medical Vision-Language Models
Abhishek Basu, Fahad Shamshad, Ashshak Sharifdeen, Karthik Nandakumar, Muhammad Haris Khan
机构
*
Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(计算机视觉系,Mohamed bin Zayed人工智能大学)
;
Department of Computer Science and Engineering, Michigan State University (MSU)(计算机科学与工程系,密歇根州立大学)