Local Margin Restoration for Test-Time Adaptation of Vision-Language Models
用于视觉-语言模型测试时适应的局部间隔恢复
Yan Huang, Guowei Wang, Xu Wang, Kangjun Liu, Xin Lin
机构
*
Guangzhou University(广州大学)
;
The Second Affiliated Hospital of Guangzhou University of Chinese Medicine(广州中医药大学第二附属医院)
;
Jinan University(暨南大学)
;
Pengcheng Laboratory(鹏城实验室)
Unifying Adversarially Robust Model Experts in Vision-Language Models
统一视觉-语言模型中的对抗鲁棒模型专家
Nguyen Duc Thai, Junhao Dong, Sua Qi Rong, Hua Yu, Yew-Soon Ong
机构
*
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
Center for Frontier AI Research, Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局前沿人工智能研究中心)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
高质量文本,稳健视觉:语言在增强视觉语言模型视觉稳健性中的作用
Futa Waseda, Saku Sugawara, Isao Echizen
机构
*
The University of Tokyo(东京大学)
;
National Institute of Informatics(日本信息处理研究所)
;
The University of Tokyo, National Institute of Informatics(东京大学、日本信息处理研究所)
机构
*
School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络空间安全学院)
;
College of Computer Science, Chongqing University(重庆大学计算机科学学院)
;
School of Software and engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)
专题命中
幻觉与鲁棒性
:vision-language model(title);vision language model(abstract);分类 cs.CV
SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
SAVER:通过风格感知视觉早期修正减轻大型视觉语言模型中的幻觉
Zhaoxu Li, Chenqi Kong, Yi Yu, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Kot, Xudong Jiang
机构
*
ROSE Lab, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(南洋理工大学罗思实验室,跨学科研究生项目,新加坡)
;
ROSE Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学罗思实验室,电子与电气工程学院,新加坡)
;
City University of Hong Kong, Hong Kong SAR(香港城市大学,香港特别行政区)
;
Shanghai Jiao Tong University, China(上海交通大学,中国)
;
Singapore University of Technology and Design, Singapore(新加坡科技设计大学,新加坡)
;
VinUniversity, Hanoi, Vietnam(越南文大学,河内,越南)
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling
面向多模态大语言模型的长视频快速有效理解:自适应准高斯采样
Kun Zhang, Chenxin Fang, Tao Chen, Baiyang Song, Yunhang Shen, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
专题命中
幻觉与鲁棒性
:multimodal large language model(title,abstract);分类 cs.CV
机构
*
University of Macau(澳门大学)
;
Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院)
;
Peking University(北京大学)
;
Independent Researcher(独立研究员)
;
Institute of Science Tokyo(东京科学研究院)
;
Morgan Stanley(摩根大通)
;
Halmstad University(哈马碧大学)
专题命中
幻觉与鲁棒性
:vision language model(title,abstract);分类 cs.CV
Mitigating Hallucinations in Large Vision-Language Models via Causal Route Gating
通过因果路由门控减轻大型视觉语言模型中的幻觉
Zhe Cheng, Wenyu Chen, Fode Zhang, Dehuan Shen
机构
*
Center of Statistical Research, School of Statistics and Data Science, Southwestern University of Finance and Economics, Chengdu, China.(统计研究中心,统计与数据科学学院,西南财经大学,成都,中国)
;
Department of Biomedical Engineering, College of Design and Engineering, National University of Singapore, Singapore(生物医学工程系,设计与工程学院,新加坡国立大学,新加坡)
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation
HalluCXR: 评估和缓解医疗视觉-语言模型在胸部X光解读中的幻觉
Haoyu Wang, Zitong Li
机构
*
Department of Biostatistics & Health Informatics, Institute of Psychiatry, Psychology & Neuroscience, King’s College London(生物统计学与健康信息学系,精神病学、心理学与神经科学研究所,伦敦国王学院)
Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models
迈向细粒度鲁棒性:面向视觉-语言模型的注意力引导测试时提示调优
Jia-Wei Hai, Yijun Wang, Xiu-Shen Wei
机构
*
School of Computer Science and Engineering(计算机科学与工程学院)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications(新一代人工智能技术及其交叉应用重点实验室)
;
Southeast University(东南大学)
;
School of Intelligence Science and Engineering(智能科学与工程学院)
机构
*
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所国家工程研究中心与人工智能院,中国科学院)
;
School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
C-CoT:基于视觉-语言模型的反事实链式推理用于安全自动驾驶
Kefei Tian, Yuansheng Lian, Kai Yang, Xiangdong Chen, Shen Li
机构
*
College of Transportation, Tongji University(同济大学交通运输学院)
;
Department of Civil Engineering, Tsinghua University(清华大学土木工程系)
;
School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)
;
Department of Civil and Environmental Engineering, National University of Singapore(新加坡国立大学土木与环境工程系)