arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-07-30 至 2025-07-30 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 7 篇

2403.15836 2025-07-30 cs.CV 88%

VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification

Lanfeng Zhong, Zongyao Huang, Yang Liu, Wenjun Liao, Shichuan Zhang, Guotai Wang, Shaoting Zhang

机构 * School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China(电子科技大学机械与电子工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Department of Pathology, Sichuan Clinical Research Center for Cancer, Sichuan Cancer Hospital & Institute, Affiliated Cancer Hospital of University of Electronic Science and Technology of China(pathology department, 四川省癌症临床研究中心, 四川省肿瘤医院及研究所, 电子科技大学附属肿瘤医院) Department of Radiation Oncology, Sichuan Cancer Hospital and Institute, University of Electronic Science and Technology of China(放射肿瘤科, 四川省肿瘤医院及研究所, 电子科技大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(title,abstract);分类 cs.CV

Comments Accepted at TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21637 2025-07-30 cs.AI 79%

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models

Wanying Wang, Zeyu Ma, Han Zheng, Xin Tan, Mingang Chen

机构 * Shanghai Key Laboratory of Computer Software Testing and Evaluating(上海软件测试与评估 key laboratory) Shanghai Normal University(上海Normal University) TrustAI Pte. Ltd. East China Normal University(东华师范大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.AI

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21521 2025-07-30 cs.CV 79%

Optimizing Active Learning in Vision-Language Models via Parameter-Efficient Uncertainty Calibration

Athmanarayanan Lakshmi Narayanan, Amrutha Machireddy, Ranganath Krishnan

机构 * Intel Labs(英特尔实验室) Intel Corporation(英特尔公司)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

Comments International Joint Conference on Neural Networks 2025 (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21494 2025-07-30 cs.LG 79%

Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated Learning

Wenxuan Bao, Ruxi Deng, Ruizhong Qiu, Tianxin Wei, Hanghang Tong, Jingrui He

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.LG

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21794 2025-07-30 cs.CV 74%

Distribution-Based Masked Medical Vision-Language Model Using Structured Reports

Shreyank N Gowda, Ruichi Zhang, Xiao Gu, Ying Weng, Lu Yang

机构 * School of Computer Science, University of Nottingham(计算机科学学院,诺丁汉大学) Department of Computer Science and Technology, School of Informatics, Xiamen University(计算机科学与技术系,信息学院,厦门大学) CHI Lab, University of Oxford(CHI实验室,牛津大学) School of Computer Science, University of Nottingham Ningbo China(计算机科学学院,宁波大学中国)

专题命中 幻觉与鲁棒性 :vision-language model(title);分类 cs.CV

Comments Accepted in MICCAI-W 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 57%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21489 2025-07-30 cs.CV 57%

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

Zhichuan Wang, Yang Zhou, Zhe Liu, Rui Yu, Song Bai, Yulong Wang, Xinwei He, Xiang Bai

机构 * Huazhong Agricultural University(华中农业大学) Shenzhen University(深圳大学) The University of Hong Kong(香港大学) University of Louisville(路易斯安那大学) ByteDance(字节跳动) Huazhong University of Science and Technology(华中科技大学)

专题命中 幻觉与鲁棒性 :MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏