arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-14 至 2025-08-14 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 7 篇

2410.19925 2025-08-14 cs.CL cs.CV cs.LG 84%

Improving Multimodal Large Language Models Using Continual Learning

Shikhar Srivastava, Md Yousuf Harun, Robik Shrestha, Christopher Kanan

机构 * University of Rochester(罗切斯特大学) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 幻觉与鲁棒性 :multimodal large language model(title);LLaVA(abstract);MLLM(abstract);分类 cs.CV、cs.LG

Comments CoLLAs 2025 and Scalable Continual Learning for Lifelong Foundation Models, NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22398 2025-08-14 cs.CV 84%

On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations

Jordan Vice, Naveed Akhtar, Yansong Gao, Richard Hartley, Ajmal Mian

机构 * University of Western Australia(西澳大学) University of Melbourne(墨尔本大学) Australian National University(澳大利亚国立大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

Comments Keywords: Vision-Language Models, Frequency-Domain Perturbations, Adversarial Robustness, Image Authenticity, Reliability

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05148 2025-08-14 cs.RO cs.AI 83%

Chemist Eye: A Visual Language Model-Powered System for Safety Monitoring and Robot Decision-Making in Self-Driving Laboratories

Francisco Munguia-Galeano, Zhengxue Zhou, Satheeshkumar Veeramani, Hatem Fakhruldeen, Louis Longley, Rob Clowes, Andrew I. Cooper

机构 * Cooper Group, Department of Chemistry, University of Liverpool(Cooper集团,化学系,利物浦大学)

专题命中 幻觉与鲁棒性 :visual language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09230 2025-08-14 cs.MA cs.AI 83%

Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems

Yutong Wu, Jie Zhang, Yiming Li, Chao Zhang, Qing Guo, Nils Lukas, Tianwei Zhang

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore, Singapore.(南洋理工大学计算机与数据科学学院) CFAR and IHPC, Agency for Science, Technology and Research, Singapore.(科技研究局) Network and Information Security Lab, Tsinghua University, Beijing, China.(清华大学网络与信息安全部门) Mohamed bin Zayed University of Artificial Intelligence, Masdar City, Abu Dhabi(马斯达尔城人工智能大学)

专题命中 幻觉与鲁棒性 :VLM(title,abstract);vision language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09218 2025-08-14 cs.CV cs.AI 81%

Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity

Zuoou Li, Weitong Zhang, Jingyuan Wang, Shuyuan Zhang, Wenjia Bai, Bernhard Kainz, Mengyun Qiao

专题命中 幻觉与鲁棒性 :MLLM(title);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09500 2025-08-14 cs.CV 79%

Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations

Yiwen Liang, Hui Chen, Yizhe Xiong, Zihan Zhou, Mengyao Lyu, Zijia Lin, Shuaicheng Niu, Sicheng Zhao, Jungong Han, Guiguang Ding

机构 * Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted at the 33rd ACM International Conference on Multimedia(ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03012 2025-08-14 cs.AI cs.CL cs.CV 62%

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

Pegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny, Matthieu Cord

机构 * ISIR, Sorbonne Université(ISIR,索邦大学)

专题命中 幻觉与鲁棒性 :MLLM(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. The first three authors contributed equally. Project page and code: https://pegah- kh.github.io/projects/lmm-finetuning-analysis-and-steering/

详情

展开后加载摘要…

URL PDF HTML 收藏