arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 2272 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 2272 篇

2509.15482 2025-11-07 cs.CV cs.AI 62%

Comparing Computational Pathology Foundation Models using Representational Similarity Analysis

Vaibhav Mishra, William Lotter

机构 * Dana-Farber Cancer Institute(达纳-法伯癌症研究所) Brigham and Women’s Hospital & Harvard Medical School(布里奇沃特医院及哈佛医学院)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Proceedings of the 5th Machine Learning for Health (ML4H) Symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26466 2025-11-04 cs.CV cs.LG 62%

Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition

Pei Peng, MingKun Xie, Hang Hao, Tong Jin, ShengJun Huang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00831 2025-11-04 cs.CV cs.AI 62%

Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack

Xin Liu, Aoyang Zhou, Aoyang Zhou

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by NAACL2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00446 2025-11-04 cs.CV cs.CR cs.LG 62%

ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training

Xin Yao, Haiyang Zhao, Yimin Chen, Jiawei Guo, Kecheng Huang, Ming Zhao

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) Miner School of Computer & Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校Miner计算机与信息科学学院)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27265 2025-11-03 cs.CV cs.LG 62%

T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis

Raza Imam, Hu Wang, Dwarikanath Mahapatra, Mohammad Yaqub

机构 * Mohammed bin Zayed University of Artificial Intelligence(莫卧儿bin Zayed人工智能大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Main: 11 pages, Supplementary: 9 pages 10 tables, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11654 2025-10-28 cs.LG cs.AI eess.SP 62%

R-SFLLM: Jamming Resilient Framework for Split Federated Learning with Large Language Models

Aladin Djuhera, Vlad C. Andrei, Xinyang Li, Ullrich J. Mönich, Holger Boche, Walid Saad

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Transactions on Information Forensics and Security (Volume: 20), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02456 2025-10-21 cs.LG cs.AI cs.NA math.NA 62%

Market-Driven Subset Selection for Budgeted Training

Ashish Jha, Valentin Leplat, AH Phan

机构 * Skolkovo Institute of Science and Technology(斯克洛夫诺科学与技术研究所) Innopolis University(因诺波利斯大学)

专题命中 幻觉与鲁棒性 :grounding(abstract);分类 cs.AI、cs.LG

Comments Retitled major revision of the same work (formerly "Market-Based Data Subset Selection -- Principled Aggregation of Multi-Criteria Example Utility"). Abstract and exposition revised; ablations added; theory clarified. Core results unchanged. Supersedes v1; please process as a replacement

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07035 2025-10-09 cs.LG cs.AI 62%

Unified Molecule Pre-training with Flexible 2D and 3D Modalities: Single and Paired Modality Integration

Tengwei Song, Min Wu, Yuan Fang

机构 * Computational Bioscience Research Center, King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹科技大学计算生物科学研究中心) Institute for Infocomm Research, A*STAR(信息与通信研究机构,A*STAR) School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算机与信息系统学院)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG

Comments CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01494 2025-10-06 cs.LG cs.AI 62%

Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed

Isha Gupta, Rylan Schaeffer, Joshua Kazdan, Ken Ziyu Liu, Sanmi Koyejo

机构 * ETH Zürich(苏黎世联邦理工学院) Stanford CS(斯坦福大学计算机科学系)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01649 2025-10-03 cs.LG cs.AI 62%

Source-Free Cross-Domain Continual Learning

Muhammad Tanzil Furqon, Mahardhika Pratama, Igor Škrjanc, Lin Liu, Habibullah Habibullah, Kutluyil Dogancay

机构 * STEM, University of South Australia(南澳大利亚大学STEM学院) Faculty of Electrical and Computer Engineering, University of Ljubljana(卢布尔雅那大学电气与计算机工程学院)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21360 2025-09-29 cs.CV cs.AI 62%

Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models

Xingkai Peng, Jun Jiang, Meng Tong, Shuai Li, Weiming Zhang, Nenghai Yu, Kejiang Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 幻觉与鲁棒性 :visual language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06336 2025-09-16 cs.CV cs.AI cs.CR 62%

Multi-View Slot Attention Using Paraphrased Texts for Face Anti-Spoofing

Jeongmin Yu, Susang Kim, Kisu Lee, Taekyoung Kwon, Won-Yong Shin, Ha Young Kim

机构 * Yonsei University(延世大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01064 2025-09-16 cs.CV cs.AI 62%

Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs

Yudong Zhang, Ruobing Xie, Yiqing Huang, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Di Wang, Yu Wang

机构 * Tsinghua University, Tencent(清华大学,腾讯) Tencent(腾讯) University of Science and Technology Beijing(北京科技大学) Tencent, University of Macau(腾讯,澳门大学) Tsinghua University(清华大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM Multimedia 2025 BNI track (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08040 2025-09-09 cs.LG cs.AI 62%

BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models

Maozhen Zhang, Mengnan Zhao, Wei Wang, Bo Wang

机构 * School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学) School of Computer Science and Technology, Anhui University(计算机科学与技术学院,安徽大学) New Laboratory of Pattern Recognition (NLPR) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS) Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室(NLPR)多模态人工智能系统国家重点实验室(MAIS)自动化研究所,中国科学院(CASIA))

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21496 2025-09-03 cs.CV cs.AI 62%

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, Xuanyu Zheng, Yepeng Tang, Dahua Lin, Lewei Lu

机构 * Sensetime(秒氏科技)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20760 2025-09-03 cs.CV cs.AI 62%

Occlusion Robustness of CLIP for Military Vehicle Classification

Jan Erik van Woerden, Gertjan Burghouts, Lotte Nijskens, Alma M. Liezenga, Sabina van Rooij, Frank Ruis, Hugo J. Kuijf

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments To be presented at SPIE: Sensors + Imaging, Artificial Intelligence for Security and Defence Applications II

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02129 2025-09-03 cs.LG cs.CV 62%

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

Jintao Cheng, Weibin Li, Jiehao Luo, Xiaoyu Tang, Zhijian He, Jin Wu, Yao Zou, Wei Zhang

机构 * Hong Kong University of Science(香港科学与技术大学) South China Normal University, Shanwei, Guangdong, China(华南师范大学,汕尾,广东,中国) Shenzhen Technology University, Shenzhen, Guangdong, China(深圳科技大学,深圳,广东,中国) University of Science(科学大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19322 2025-08-28 eess.IV cs.AI cs.CV 62%

AT-CXR: Uncertainty-Aware Agentic Triage for Chest X-rays

Xueyang Li, Mingze Jiang, Gelei Xu, Jun Xia, Mengzhao Jia, Danny Chen, Yiyu Shi

机构 * University of Notre Dame(诺丁汉大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11341 2025-08-18 cs.CV cs.CR cs.LG 62%

Semantically Guided Adversarial Testing of Vision Models Using Language Models

Katarzyna Filus, Jorge M. Cruz-Duarte

机构 * Institute of Theoretical and Applied Informatics, Polish Academy of Sciences(波兰科学院理论与应用信息学研究所) University of Lille, CNRS, Inria, Centrale Lille, UMR 9189 CRIStAL(里尔大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

Comments 12 pages, 4 figures, 3 tables. Submitted for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03012 2025-08-14 cs.AI cs.CL cs.CV 62%

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

Pegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny, Matthieu Cord

机构 * ISIR, Sorbonne Université(ISIR,索邦大学)

专题命中 幻觉与鲁棒性 :MLLM(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. The first three authors contributed equally. Project page and code: https://pegah- kh.github.io/projects/lmm-finetuning-analysis-and-steering/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03164 2025-08-06 cs.CV cs.AI cs.CL 62%

ChartCap: Mitigating Hallucination of Dense Chart Captioning

Junyoung Lim, Jaewoo Ahn, Gunhee Kim

机构 * Seoul National University(首尔国立大学)

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20028 2025-07-29 cs.CV cs.AI 62%

TAPS : Frustratingly Simple Test Time Active Learning for VLMs

Dhruv Sarkar, Aprameyo Chakrabartty, Bibhudatta Bhanja

机构 * IIT Kharagpur(印度理工学院Kharagpur分校)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15265 2025-07-28 cs.CV cs.AI cs.CR 62%

Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs

Zihao Pan, Yu Tong, Weibin Wu, Jingyi Wang, Lifeng Chen, Zhe Zhao, Jiajia Wei, Yitong Qiao, Zibin Zheng

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments The paper needs major revisions, so it is being withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09222 2025-07-22 cs.CV cs.LG 62%

Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift

Behraj Khan, Tahir Qasim Syed, Nouman M. Durrani, Bilal Naseem, Shabir Ahmad, Rizwan Qureshi

机构 * Institute of Business Administration Karachi(Karachi商业管理学院) National University of Computer and Emerging Sciences(国家计算机与新兴科学大学) CAIMI Pvt Ltd(CAIMI私营有限公司) Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,佛罗里达大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05020 2025-07-11 cs.CV cs.AI 62%

Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision

Soham Walimbe, Britty Baby, Vinkle Srivastav, Nicolas Padoy

机构 * University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学,法国国家科学研究中心(CNRS),法国国家卫生研究院(INSERM),ICube,UMR7357,斯特拉斯堡) Institute of Image-Guided Surgery, IHU Strasbourg, Strasbourg, France(影像引导手术研究所,斯特拉斯堡IHU,斯特拉斯堡)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03589 2025-07-08 cs.CV cs.AI cs.CL 62%

BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance

Huy Le, Nhat Chung, Tung Kieu, Anh Nguyen, Ngan Le

机构 * FPT Software AI Center(FPT软件AI中心) Aalborg University(奥尔堡大学) Pioneer Centre for AI(先锋人工智能中心) University of Liverpool(利物浦大学) AICV Lab, University of Arkansas(AICV实验室,阿肯色大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03458 2025-07-08 cs.CV cs.AI 62%

Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach

Leyan Xue, Zongbo Han, Guangyu Wang, Qinghua Hu, Mingyue Cheng, Changqing Zhang

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01980 2025-06-30 cs.LG cs.AI 62%

Generative Data Mining with Longtail-Guided Diffusion

David S. Hayden, Mao Ye, Timur Garipov, Gregory P. Meyer, Carl Vondrick, Zhao Chen, Yuning Chai, Eric Wolff, Siddhartha S. Srinivasa

机构 * OpenAI Colombia University(哥伦比亚大学) Meta

专题命中 幻觉与鲁棒性 :VLM(abstract);分类 cs.AI、cs.LG

Comments 20 pages

Journal ref Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06967 2025-06-23 cs.CV cs.AI eess.IV 62%

Dual Thinking and Logical Processing -- Are Multi-modal Large Language Models Closing the Gap with Human Vision ?

Kailas Dayanandan, Nikhil Kumar, Anand Sinha, Brejesh Lall

机构 * Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 幻觉与鲁棒性 :vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06745 2025-06-18 cs.LG cs.CL cs.CV 62%

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Adhiraj Ghosh, Sebastian Dziadzio, Ameya Prabhu, Vishaal Udandarao, Samuel Albanie, Matthias Bethge

机构 * Tübingen AI Centre(图宾根人工智能中心) University of Tübingen(图宾根大学) University of Cambridge(剑桥大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏