arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-24 至 2025-09-24 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 7 篇

2507.13255 2025-09-24 cs.CL cs.AI cs.IR cs.LG cs.MM 84%

Automating Steering for Safe Multimodal Large Language Models

Lyucheng Wu, Mengru Wang, Ziwen Xu, Tri Cao, Nay Oo, Bryan Hooi, Shumin Deng

机构 * Zhejiang University(浙江大学) Zhejiang University - Ant Group Joint Lab of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室) National University of Singapore, NUS-NCS Joint Lab(新加坡国立大学NUS-NCS联合实验室)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);LLaVA(abstract);分类 cs.AI、cs.LG

Comments EMNLP 2025 Main Conference. 23 pages (8+ for main); 25 figures; 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19212 2025-09-24 cs.CL cs.AI 83%

Steering Multimodal Large Language Models Decoding for Context-Aware Safety

Zheyuan Liu, Zhangchen Xu, Guangyao Dou, Xiangchi Yuan, Zhaoxuan Tan, Radha Poovendran, Meng Jiang

机构 * University of Notre Dame(notre dame 大学) University of Washington(华盛顿大学) Johns Hopkins University(约翰霍普金斯大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI

Comments A lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19070 2025-09-24 cs.CV cs.CL 79%

ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?

Zijian Ling, Han Zhang, Yazhuo Zhou, Jiahao Cui

机构 * Apply U United Kingdom(英国Apply大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted at the Open Science for Foundation Models (SCI-FM) Workshop at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15389 2025-09-24 cs.CL cs.CR cs.CV 79%

Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study

DongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo Yu

机构 * Pohang University of Science and Technology (POSTECH)(釜山科学技术大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14257 2025-09-24 cs.CV 79%

Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow

Jianfei Zhao, Feng Zhang, Xin Sun, Chong Feng

机构 * School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) Zhongguancun Academy(中关村学院) Southeast Academy of Information Technology, Beijing Institute of Technology(信息技术东南学院,北京理工大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19203 2025-09-24 cs.CV 57%

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

机构 * Queen Mary University of London(伦敦女王大学) Samsung AI Centre(三星人工智能中心) Technical University of Iași(伊阿苏技术大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18998 2025-09-24 math.AP cs.LG q-bio.CB q-bio.QM 57%

Bayesian Calibration and Model Assessment of Cell Migration Dynamics with Surrogate Model Integration

Christina Schenk, Jacobo Ayensa Jiménez, Ignacio Romero

机构 * IMDEA Materials Institute(IMDEA材料研究所) Universidad Politécnica de Madrid(马德里理工大学) Aragón Institute of Engineering Research (I3A)(阿伦工程技术研究所(I3A))

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.LG

Comments 31 pages, 13 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏