Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift
Behraj Khan, Tahir Qasim Syed, Nouman M. Durrani, Bilal Naseem, Shabir Ahmad, Rizwan Qureshi
机构
*
Institute of Business Administration Karachi(Karachi商业管理学院)
;
National University of Computer and Emerging Sciences(国家计算机与新兴科学大学)
;
CAIMI Pvt Ltd(CAIMI私营有限公司)
;
Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,佛罗里达大学)
机构
*
State Key Laboratory of Internet of Things for Smart City(物联网智能城市国家重点实验室)
;
University of Macau(澳门大学)
;
Department of Civil Engineering(土木工程系)
;
Department of Computer and Information Science(计算机与信息科学系)
;
College of Transportation Engineering(交通工程学院)
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
Yi Nian, Shenzhe Zhu, Yuehan Qin, Li Li, Ziyi Wang, Chaowei Xiao, Yue Zhao
机构
*
University of Southern California(南加州大学)
;
University of Toronto(多伦多大学)
;
University of Maryland(马里兰大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion
Yuan Bian, Min Liu, Yunqi Yi, Xueping Wang, Yaonan Wang
机构
*
School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院)
;
National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心)
;
College of Information Science and Engineering, Hunan Normal University(湖南师范大学信息科学与工程学院)
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
Jie Yang, Wang Zeng, Sheng Jin, Lumin Xu, Wentao Liu, Chen Qian, Zhen Li, Ruimao Zhang
机构
*
Sun Yat-sen University, Shenzhen(中山大学深圳校区)
;
Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区)
;
SenseTime Research(商汤科技研究院)
;
Chinese University of Hong Kong(香港中文大学)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)
专题命中
图文多模态
:multimodal(abstract);分类 cs.CV
CommentsExtended Version of KptLLM. arXiv admin note: text overlap with arXiv:2411.01846
Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
Shiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu, Siyang Gao, Miao Xiong, Manling Li, Junxian He
机构
*
City University of Hong Kong(香港城市大学)
;
Hong Kong University of Science(香港科学大学)
;
National University of Singapore(新加坡国立大学)
;
Northwestern University(西北大学)
机构
*
Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)
;
Media Analytics and Computing Lab, Department of Artificial Intelligence, School of Informatics, Xiamen University(媒体分析与计算实验室,人工智能系,厦门大学)
;
School of Computer Science and Technology, Ocean University of China(计算机科学与技术学院,中国海洋大学)
;
Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(图像通信与信息处理研究所,上海交通大学)
机构
*
Department of Computer Science, University of California, Davis, CA, USA(加州大学戴维斯分校计算机科学系)
;
Bosch Center for Artificial Intelligence (BCAI), Bosch Research North America(博世人工智能中心(BCAI)、博世北美研究部)
;
Splunk Technology, San Jose, CA, USA(Splunk技术公司)
专题命中
图文多模态
:multi-modal(abstract);分类 cs.CV
CommentsIEEE Transactions on Visualization and Computer Graphics (2025)
Balancing Performance and Efficiency in Zero-shot Robotic Navigation
Dmytro Kuzmenko, Nadiya Shvai
机构
*
Department of Multimedia Systems, National University of Kyiv-Mohyla Academy, Kyiv, Ukraine(多媒体系统系,基辅-莫希拉学院国家大学,乌克兰基辅)
;
Department of Mathematics, National University of Kyiv-Mohyla Academy, Kyiv, Ukraine(数学系,基辅-莫希拉学院国家大学,乌克兰基辅)
专题命中
图文多模态
:multi-modal(abstract);分类 cs.CV
CommentsSubmitted to ICTERI 2024 Posters Track
Journal refICTERI 2024: Communications in Computer and Information Science, vol. 2020, pp. 370-381, Springer, 2025
Shih-Han Chou, Shivam Chandhok, James J. Little, Leonid Sigal
机构
*
Department of Computer Science, University of British Columbia(不列颠哥伦比亚大学计算机科学系)
;
Vector Institute for AI(人工智能矢量研究所)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)