arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4682 篇

2603.14185 2026-03-26 cs.AI 83%

Relationship-Aware Safety Unlearning for Multimodal LLMs

关系感知的安全反向学习用于多模态大语言模型

Vishnu Narayanan Anilkumar, Abhijith Sreesylesh Babu, Trieu Hai Vo, Mohankrishna Kolla, Alexander Cuneo

机构 * Florida International University(佛罗里达国际大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.AI

AI总结 本文提出关系感知的安全反向学习框架,通过显式表示不安全的对象-关系-对象元组,并应用参数高效编辑技术抑制不安全元组,同时保留对象边缘和安全邻近关系。

Comments 9 pages,4figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04559 2026-03-24 cs.CV 83%

Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning

面向可扩展多模态推理的推理对齐感知解耦

Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Xin Jin, Zhenguo Li, James T. Kwok, Yu Zhang

机构 * Southern University of Science and Technology(南方科技大学) The Hong Kong University of Science and Technology(香港科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Huawei Cloud Project(华为云项目)

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出RAPID方法,通过解耦多模态模型的感知与推理模块,利用强化学习优化感知输出,使模型能与任意强大文本推理模型配合,实现无需重新训练的性能提升。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13884 2026-03-17 cs.CV 83%

SCoCCA: Multi-modal Sparse Concept Decomposition via Canonical Correlation Analysis

SCoCCA: 通过典型相关分析进行多模态稀疏概念分解

Ehud Gordon, Meir Yossef Levi, Guy Gilboa

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出SCoCCA,通过典型相关分析实现多模态稀疏概念分解,提升跨模态对齐与可解释性,实现概念发现的最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25896 2026-03-11 cs.CV 83%

LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models

LLaVAShield: 保障视觉语言模型中的多模态多轮对话安全

Guolei Huang, Qinzhi Peng, Gan Xu, Yao Huang, Yuxuan Lu, Yongjun Shen

机构 * Southeast University(东南大学) University of California, Santa Cruz(加州大学圣克鲁兹分校) Zhejiang University of Technology(浙江工业大学) Tsinghua University(清华大学) RealAI

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 LLaVAShield通过构建MMDS数据集和MMRT框架,有效提升多模态多轮对话的安全性,优于现有VLMs和内容审查工具,揭示了主流模型的安全漏洞。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01555 2026-03-11 eess.IV cs.CV 83%

MGCR-Net:Multimodal Graph-Conditioned Vision-Language Reconstruction Network for Remote Sensing Change Detection

MGCR-Net:多模态图条件视觉-语言重建网络用于遥感变化检测

Chengming Wang, Guodong Fan, Jinjiang Li, Min Gan, C. L. Philip Chen

机构 * School of Computer Science and Technology, Shandong Technology and Business University(山东科技职业大学计算机科学与技术学院) School of Computer Science and Technology, Qingdao University(青岛大学计算机科学与技术学院) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 MGCR-Net通过多模态图条件视觉-语言重建机制提升遥感变化检测的语义交互能力。

Journal ref IEEE Transactions on Geoscience and Remote Sensing, vol. 64, pp. 1-15, 2026, Art no. 4701515

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01416 2026-03-03 cs.AI 83%

Securing the Floor and Raising the Ceiling: A Merging-based Paradigm for Multi-modal Search Agents

保障地板,提升天花板:一种基于合并的多模态搜索代理范式

Zhixiang Wang, Jingxuan Xu, Dajun Chen, Yunfang Wu, Wei Jiang, Yong Li

机构 * Peking University, Beijing, China(北京大学)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种无需训练的多模态搜索代理范式,通过跨模态模型合并提升VLMs的自主搜索能力,OBM算法在搜索任务中表现出更高的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01783 2026-03-03 cs.CV 83%

Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing

在多模态大语言模型中利用链式推理进行人脸反伪装

Honglu Zhang, Zhiqin Fang, Ningning Zhao, Saihui Hou, Long Ma, Renwang Pei, Zhaofeng He

机构 * Didi Chuxing(滴滴出行) Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出FaceCoT数据集和CEPL策略,通过链式推理提升多模态大语言模型在人脸反伪装中的鲁棒性和性能。

Comments Accepted to CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20767 2026-02-25 cs.MM 83%

SPP-SCL: Semi-Push-Pull Supervised Contrastive Learning for Image-Text Sentiment Analysis and Beyond

SPP-SCL:半推拉监督对比学习用于图像-文本情感分析及更广泛的领域

Jiesheng Wu, Shengrong Li

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 SPP-SCL通过半推拉监督对比学习平衡图像-文本情感关系,提升跨模态情感分析的性能。

Comments Accepted and published by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02669 2026-02-19 cs.CV 83%

MedVLThinker: Simple Baselines for Multimodal Medical Reasoning

MedVLThinker: 多模态医疗推理的简单基线

Xiaoke Huang, Juncheng Wu, Hui Liu, Xianfeng Tang, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) Amazon Research(亚马逊研究院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 MedVLThinker通过简单基线和RLVR方法,在医疗多模态推理中实现新突破,超越现有开源模型并接近专有模型性能。

Comments Project page: https://ucsc-vlaa.github.io/MedVLThinker/ ; Code: https://github.com/UCSC-VLAA/MedVLThinker ; Model and Data: https://huggingface.co/collections/UCSC-VLAA/medvlthinker-688f52224fb7ff7d965d581d ; Accepted by ML4H'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20322 2026-02-05 cs.CV 83%

Dynamic Pyramid Network for Efficient Multimodal Large Language Model

动态金字塔网络用于高效多模态大语言模型

Hao Ai, Kunyi Wang, Zezhou Wang, Hao Lu, Jin Tian, Yaxin Luo, Peng Xing, Jen-Yuan Huang, Huaxia Li, Gen luo

机构 * Beihang University(北航大学) Shanghai AI Laboratory(上海人工智能实验室) KAUST(卡塔尔大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Technical University Of Denmark(丹麦技术大学) Nanjing University of Science and Technology(南京理工大学) Peking University(北京大学) Xiaohongshu Inc(小红书公司)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 动态金字塔网络通过分层结构和动态池化专家提升多模态大语言模型的效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02175 2026-02-04 cs.CV 83%

CIEC: Coupling Implicit and Explicit Cues for Multimodal Weakly Supervised Manipulation Localization

CIEC: 通过隐式和显式线索耦合实现多模态弱监督操作定位

Xinquan Yu, Wei Lu, Xiangyang Luo, Rui Yang

机构 * School of Computer Science and Engineering(计算机科学与工程学院) MoE Key Laboratory of Information Technology(信息技术关键实验室) Key Laboratory of Information Security Technology(信息安全技术重点实验室) Sun Yat-sen University(中山大学) State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室) Alibaba Group(阿里巴巴集团)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 CIEC通过耦合隐式和显式线索,实现多模态弱监督操作定位,利用粗粒度标注提升图像-文本对的定位效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00135 2026-02-03 cs.CV 83%

LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models

LLaVA-FA: 通过傅里叶近似压缩大型多模态模型

Pengcheng Zheng, Chaoning Zhang, Jiarong Mo, GuoHui Li, Jiaquan Zhang, Jiahao Zhang, Sihan Cao, Sheng Zheng, Caiyan Qin, Guoqing Wang, Yang Yang

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 LLaVA-FA通过频域联合低秩和量化近似,实现高效压缩大型多模态模型,采用极坐标量化和对角校准方案,提升压缩效果与效率。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05895 2026-01-28 cs.CV 83%

BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model

BTCChat: 通过多模态大语言模型推进遥感双时相变化描述

Yujie Li, Wenjia Xu, Yuanben Zhang, Zhiwei Wei, Mugen Peng

机构 * State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室) Beijing University of Posts and Telecommunications(北京邮电大学) Aerospace Information Research Institute(航天信息研究所) Chinese Academy of Sciences(中国科学院) School of Geographical Sciences(地理科学学院) Hunan Normal University(湖南师范大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 BTCChat通过多模态大语言模型提升遥感双时相变化描述能力,引入变化提取模块和提示增强机制,实现更精确的视觉-语义对齐和更优的性能表现。

Comments 5 pages, 2 figures; Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11808 2026-01-27 cs.CV cs.AI cs.CL cs.CY cs.MM 83%

Labels or Input? Rethinking Augmentation in Multimodal Hate Detection

标签还是输入?重新思考多模态仇恨检测中的增强

Sahajpreet Singh, Kokil Jaidka, Subhayan Mukerjee

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出通过提示优化、微调和自动化数据增强改进小型模型,开发多模态增强框架以提升隐含仇恨检测性能。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23243 2026-01-12 cs.CV 83%

Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism

遥感图像的多模态解释:动态分辨率输入策略与多尺度视觉-语言对齐机制

Siyu Zhang, Lianlei Shan, Runhe Qiu

机构 * Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出动态分辨率输入策略与多尺度视觉-语言对齐机制,提升遥感图像多模态解释的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10969 2026-01-01 cs.CV 83%

Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation

弥合一致性差距:用于交错图像-文本生成的显式结构化记忆

Zeteng Lin, Xingxing Li, Wen You, Xiaoyang Li, Zehan Lu, Yujun Cai, Jing Tang

机构 * Hong Kong University of Science and Technology(Guangzhou)(香港科技大学(广州)) University of Queensland(昆士兰大学)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 IUT-Plug通过显式结构化记忆机制解决多模态生成中的上下文漂移问题,提升长序列一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18033 2025-12-29 cs.RO cs.AI 83%

OpenNav: Open-World Navigation with Multimodal Large Language Models

OpenNav: 基于多模态大语言模型的开放世界导航

Mingfeng Yuan, Letian Wang, Steven L. Waslander

机构 * University of Toronto Institute for Aerospace Studies(多伦多大学航空航天研究 institute) University of Toronto Robotics Institute(多伦多大学机器人研究所)

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 OpenNav利用多模态大语言模型实现开放世界导航,通过生成价值图增强机器人空间理解,并在真实场景中验证其鲁棒性。

Journal ref 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21583 2025-12-29 cs.AI 83%

A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning

一种整合视觉-语言模型和逻辑树推理的医学多模态诊断框架

Zelin Zang, Wenyi Gu, Siqi Ma, Dan Yang, Yue Shen, Zhu Zhang, Guohui Fan, Wing-Kuen Ling, Fuji Yang

机构 * Tsientang Institute of Advanced Study (TIAS)(钱塘先进研究所) Westlake University(西湖大学) Ant Group(蚂蚁集团) China-Japan Friendship Hospital(中日友好医院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出一种整合视觉-语言模型与逻辑树推理的医学诊断框架,旨在提升多模态医学AI的可信度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20257 2025-12-24 cs.CV 83%

LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation

LADLE-MM:基于有限标注的多模态虚假信息检测器

Daniele Cardullo, Simone Teglia, Irene Amerini

机构 * Sapienza University of Rome(罗马萨皮恩扎大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 LADLE-MM是一种基于有限标注的多模态虚假信息检测器,通过学习的集成方法在有限资源下实现高效检测,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17178 2025-12-22 cs.CV cs.IR 83%

ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching

ABE-CLIP: 无训练属性绑定增强用于组合图像-文本匹配

Qi Zhang, Yuxu Chen, Lei Deng, Lili Shen

机构 * School of Mathematics, Sichuan University(四川大学数学学院)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 ABE-CLIP通过语义细化和局部对齐策略提升CLIP模型的属性-物体绑定性能,无需额外训练。

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12107 2025-12-16 cs.CV 83%

EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography

EchoVLM:基于测量的多模态学习用于超声心动图

Yuheng Li, Yue Zhang, Abdoul Aziz Amadou, Yuxiang Lai, Jike Zhong, Tiziano Passerini, Dorin Comaniciu, Puneet Sharma

机构 * Georgia Institute of Technology(佐治亚理工学院) Siemens Healthineers(西门子医疗) Siemens Healthcare Limited(西门子医疗有限公司) Emory University(埃默里大学) University of Southern California(南加州大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 EchoVLM通过引入基于测量的多模态学习方法,实现了超声心动图的端到端解读,提升了疾病分类和视图识别的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02278 2025-12-09 cs.CV 83%

SCMM: Calibrating Cross-modal Representations for Text-Based Person Search

SCMM:基于文本的人脸搜索的跨模态表示校准

Jing Liu, Donglai Wei, Yang Liu, Sipeng Zhang, Tong Yang, Wei Zhou, Weiping Ding, Victor C. M. Leung

机构 * College of Future Information Technology, Fudan University(复旦大学未来信息技术学院) Department of Electrical and Computer Engineering, The University of British Columbia(不列颠哥伦比亚大学电气与计算机工程系) College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院) MEGVII Technology(MEGVII技术) School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) School of AI and CS, Nantong University(南通大学人工智能与计算机科学学院) Academy of Artificial Intelligence, SMBU(SMBU人工智能学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 SCMM通过缝校准和掩码建模方法,提升跨模态表示学习在文本驱动的人脸搜索中的性能。

Comments 11 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00396 2025-11-27 cs.CV 83%

Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning

Saliency-R1: 通过置信度引导的强化学习激励多模态大语言模型统一的显著性推理能力

Long Li, Shuichen Ji, Ziyang Luo, Zhihui Li, Dingwen Zhang, Junwei Han, Nian Liu

机构 * Northwestern Polytechnical University(西北工业大学) University of Science and Technology of China(中国科学技术大学)

专题命中 图文多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 Saliency-R1通过置信度引导的强化学习,提升多模态大语言模型在显著性推理任务中的统一能力,实现显著物体检测、显著实例分割和共显著物体检测的高效处理。

Comments Main text (excluding references): 8 pages, 4 figures; Supplementary Materials (excluding references): 9 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12215 2025-11-18 cs.CV 83%

FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention

Peng Zhang, Zhihui Lai, Wenting Chen, Xu Wu, Heng Kong

机构 * Peng Zhang 1 , Zhihui Lai 1 , Wenting Chen 2 1 1 footnotemark: 1 , Xu Wu 1 , Heng Kong 3(某机构)

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11984 2025-11-18 cs.CV 83%

From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology

Zhenhao Guo, Rachit Saluja, Tianyuan Yao, Quan Liu, Junchao Zhu, Haibo Wang, Daniel Reisenbüchler, Yuankai Huo, Benjamin Liechty, David J. Pisapia, Kenji Ikemura, Steven Salvatoree, Surya Seshane, Mert R. Sabuncu, Yihe Yang, Ruining Deng

机构 * New York University(纽约大学) Cornell Tech(康奈尔科技) Vanderbilt University(范德比大学) Carnegie Mellon University(卡内基梅隆大学) University of Regensburg(莱茵河畔大学) Weill Cornell Medicine(韦尔·科恩医学中心) Northwell Health(北well健康)

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21375 2025-11-05 cs.CV 83%

GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution

Fengxiang Wang, Mingshuo Chen, Yueying Li, Di Wang, Haotian Wang, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang, Hongzhen Wang, Wenjing Yang, Bo Du, Jing Zhang

机构 * College of Computer Science and Technology, National University of Defense Technology, China(中国国防科技大学计算机科学与技术学院) Beijing University of Posts and Telecommunications, China(北京邮电大学) School of Computer Science, Wuhan University, China(武汉大学计算机学院) Zhongguancun Academy, China(中关村学院) Tsinghua University, China(清华大学) Beihang University, China(北航大学)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments NeurlPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20392 2025-10-31 cs.CV 83%

Defending Multimodal Backdoored Models by Repulsive Visual Prompt Tuning

Zhifang Zhang, Shuo He, Haobo Wang, Bingquan Shen, Lei Feng

机构 * Southeast University(东南大学) University of Queensland(昆士兰大学) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25303 2025-10-30 cs.CL 83%

Teaching Sarcasm: Few-Shot Multimodal Sarcasm Detection via Distillation to a Parameter-Efficient Student

Soumyadeep Jana, Sanasam Ranbir Singh

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Indian Institute of Technology Guwahati(印度理工学院古瓦哈蒂)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13497 2025-10-16 cs.LG cs.AI 83%

DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation

Zexin Wang, Lin Shi, Haoyu Wu, Junru Luo, Xiangzeng Kong, Jun Qi

机构 * Aliyun School of Big Data, Changzhou University(阿里云大数据学院,长洲大学) Department of Computing, Xi’an JiaoTong-Liverpool University(计算系,西安交通大学-利物浦大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Center for Artificial Intelligence in Agriculture, Fujian Agriculture and Forestry University(农业人工智能中心,福建农林大学)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10466 2025-10-14 cs.CV 83%

When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance

Jinjin Cao, Zhiyang Chen, Zijun Wang, Liyuan Ma, Weijian Luo, Guojun Qi

机构 * MAPLE Lab, Westlake University(西溪大学MAPLE实验室)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏