arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46237 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4672 篇

2511.22906 2025-12-01 cs.CV 70%

See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection

见、排、滤:通过场景理解的重要词感知Clip过滤用于片段检索和亮点检测

YuEun Lee, Jung Uk Kim

机构 * YuEun Lee, Jung Uk Kim

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出通过识别查询中的重要词,结合多模态大语言模型实现视频片段检索和亮点检测的细粒度过滤方法,提升检索和检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22256 2025-12-01 cs.CV 70%

UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation

UMind-VL:一种通用的超声视觉-语言模型,用于统一的 grounded perception 和全面的 interpretation

Dengbo Chen, Ziwei Zhao, Kexin Zhang, Shishuang Zhao, Junjie Hou, Yaqian Wang, Nianxi Liao, Anlan Sun, Fei Gao, Jia Ding, Yuhang Liu, Dong Wang

机构 * Yizhun Medical AI Team(义诊医疗AI团队)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 UMind-VL 是一种通用超声视觉-语言模型,通过统一的 grounded perception 和 comprehensive interpretation 实现对医学影像的高效理解和诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19516 2025-11-27 cs.CV 70%

Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning

连接点:通过代理推理实现无需训练的视觉接地

Liqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai, Yixiong Zou, Yonghong Tian

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 GroundingAgent通过代理推理实现无需微调的视觉接地,达到65.1%的零样本准确率,并在选择阶段实现90%的准确率,展示了LLM推理能力的重要性。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16181 2025-11-26 cs.CV 70%

CLIP-IT: CLIP-based Pairing for Histology Images Classification

CLIP-IT: 基于CLIP的病理图像分类配对

Banafsheh Karimian, Giulia Avanzato, Soufian Belharbi, Alexis Guichemerre, Luke McCaffrey, Mohammadhadi Shateri, Eric Granger

机构 * LIVIA, ILLS, Dept. of Systems Engineering, ETS Montreal, Canada(LIVIA、ILLs、系统工程系、蒙特利尔大学ETSMontreal加拿大) Dept. of Computer Engineering, University of Cagliari, Italy(计算机工程系、卡利亚里大学意大利) Goodman Cancer Research, Centre, Dept. of Oncology, McGill University, Canada(Goodman癌症研究中心、肿瘤学系、麦吉尔大学加拿大)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 CLIP-IT通过利用未配对的病理报告提升病理图像分类性能,无需配对数据或复杂推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18921 2025-11-25 cs.CV 70%

BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models

BackdoorVLM:面向视觉-语言模型的后门攻击基准

Juncheng Li, Yige Li, Hanxun Huang, Yunhao Chen, Xin Wang, Yixu Wang, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Singapore Management University(新加坡管理学院) The University of Melbourne(墨尔本大学)

专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 BackdoorVLM提出首个评估视觉-语言模型后门攻击的基准,揭示VLMs对文本指令的高敏感性及多模后门的有效性

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17793 2025-11-25 cs.CV cs.LG 70%

Attention Guided Alignment in Efficient Vision-Language Models

注意力引导的高效视觉-语言模型

Shweta Mahajan, Hoang Le, Hyojin Park, Farzad Farhadzadeh, Munawar Hayat, Fatih Porikli

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出AGE-VLM,通过交错交叉注意力层和空间知识提取,减少高效视觉-语言模型中的幻觉问题。

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13889 2025-11-20 cs.CV cs.LG 70%

Uni-Hema: Unified Model for Digital Hematopathology

Abdul Rehman, Iqra Rasool, Ayisha Imran, Mohsen Ali, Waqas Sultani

机构 * Information Technology University of Punjab(旁遮普信息科技大学) Chughtai Lab(楚格塔实验室)

专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13876 2025-11-19 cs.CV 70%

QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning

Xiaoyang Wei, Camille Kurtz, Florence Cloppet

机构 * Laboratoire d'Informatique Paris Descartes (LIPADE), Université Paris Cité (France)(巴黎笛卡尔大学信息学实验室(LIPADE),巴黎城市大学(法国))

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE ISBI for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02329 2025-11-18 cs.CV 70%

VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions

Ziteng Wang, Siqi Yang, Limeng Qiao, Lin Ma

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17417 2025-11-13 cs.CV 70%

Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment

Robert Wijaya, Ngoc-Bao Nguyen, Ngai-Man Cheung

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05057 2025-11-10 cs.CV 70%

Role-SynthCLIP: A Role Play Driven Diverse Synthetic Data Approach

Yuanxiang Huangfu, Chaochao Wang, Weilei Wang

机构 * PatSnap Co., LTD.(PatSnap公司)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21849 2025-11-07 cs.LG cs.AI 70%

TowerVision: Understanding and Improving Multilinguality in Vision-Language Models

André G. Viveiros, Patrick Fernandes, Saul Santos, Sonal Sannigrahi, Emmanouil Zaranis, Nuno M. Guerreiro, Amin Farajian, Pierre Colombo, Graham Neubig, André F. T. Martins

机构 * Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术高级学院) Instituto de Telecomunicações(电信研究所) Carnegie Mellon University(卡内基梅隆大学) Sword Health(Sword健康) TransPerfect MICS, CentraleSupélec, Université Paris-Saclay(MICS,中央圣艾尔布里大学,巴黎萨克雷大学) ELLIS Unit Lisbon(里斯本ELLIS单位)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.AI

Comments 15 pages, 7 figures, submitted to arXiv October 2025. All models, datasets, and training code will be released at https://huggingface.co/collections/utter-project/towervision

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04220 2025-11-06 cs.CV 70%

Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs

Fangrui Zhu, Hanhui Wang, Yiming Xie, Jing Gu, Tianye Ding, Jianwei Yang, Huaizu Jiang

机构 * Northeastern University(东北大学) Microsoft Research(微软研究院) University of Southern California(南加州大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS 2025, code link: https://github.com/neu-vi/struct2d

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25070 2025-10-30 cs.CV 70%

Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments

Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22785 2025-10-28 cs.CV 70%

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari, Mingkun Xu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Agency for Science, Technology and Research(科技研究局) School of Information Technology(信息技术学院)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07019 2025-10-28 cs.CV 70%

A Vision-Language Foundation Model for Leaf Disease Identification

Khang Nguyen Quoc, Lan Le Thi Thu, Luyl-Da Quach

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21606 2025-10-27 cs.CV 70%

Modest-Align: Data-Efficient Alignment for Vision-Language Models

Jiaxiang Liu, Yuan Wang, Jiawei Du, Joey Tianyi Zhou, Mingkun Xu, Zuozhu Liu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) ZJU-Angelalign R&D Center for Intelligence Healthcare(浙大天使align智能医疗研发中心) Centre for Frontier AI Research (CFAR)(前沿人工智能研究中心) Agency for Science, Technology and Research (A*STAR)(科技研究局) Institute of High Performance Computing (IHPC)(高性能计算研究所)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21311 2025-10-27 cs.CV 70%

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

Lu Zhang, Jiazuo Yu, Haomiao Xiong, Ping Hu, Yunzhi Zhuge, Huchuan Lu, You He

机构 * Dalian University of Technology(大连理工大学) University of Electronic Science and Technology of China(电子科技大学) Tsinghua Shenzhen International Graduate School(清华大学深圳国际graduate school)

专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16870 2025-10-21 cs.CV 70%

Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding

Yudan Ren, Xinlong Wang, Kexin Wang, Tian Xia, Zihan Ma, Zhaowei Li, Xiangrong Bi, Xiao Li, Xiaowei He

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13946 2025-10-21 cs.AI 70%

Visual Instruction Bottleneck Tuning

Changdae Oh, Jiatong Li, Shawn Im, Sharon Li

机构 * Department of Computer Sciences, University of Wisconsin–Madison(计算机科学系,威斯康星大学麦迪逊分校)

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13235 2025-10-16 cs.CV 70%

EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking

Yukuan Zhang, Jiarui Zhao, Shangqing Nie, Jin Kuang, Shengsheng Wang

机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(吉林大学教育部长春符号计算与知识工程重点实验室) Yangtze University(扬子大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01738 2025-10-14 cs.CV 70%

DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy

Ming Dai, Wenxuan Cheng, Jiang-jiang Liu, Sen Yang, Wenxiao Cai, Yanpeng Sun, Wankou Yang

机构 * Southeast University(东南大学) Baidu VIS(百度VIS) Stanford University(斯坦福大学)

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15130 2025-10-10 cs.CV 70%

TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba

Xiuwei Chen, Wentao Hu, Xiao Dong, Sihao Lin, Zisheng Chen, Meng Cao, Yina Zhuang, Jianhua Han, Hang Xu, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Hong Kong Polytechnic University(香港理工大学) Zhuhai Campus of Sun Yat-sen University(中山大学珠海校区) University of Adelaide(阿德莱德大学) Peking University(北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02815 2025-10-06 cs.CV 70%

Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis

Feng Yuan, Yifan Gao, Yuehua Ye, Haoyue Li, Xin Gao

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute of Biomedical Engineering and Technology(苏州生物医学工程与技术研究所) Chinese Academy of Sciences(中国科学院) The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院)

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments ICLR2026 under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25654 2025-10-01 cs.CV 70%

DescribeEarth: Describe Anything for Remote Sensing Images

Kaiyu Li, Zixuan Jiang, Xiangyong Cao, Jiayu Wang, Yuchen Xiao, Deyu Meng, Zhi Wang

机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室) School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室) Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)

专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06105 2025-10-01 cs.CV 70%

PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology

Yating Huang, Ziyan Huang, Lintao Xiang, Qijun Yang, Hujun Yin

机构 * University of Manchester(曼彻斯特大学) South China University of Technology(华南理工大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accept by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24258 2025-09-30 cs.CV 70%

When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs

Jinming Liu, Zhaoyang Jia, Jiahao Li, Bin Li, Xin Jin, Wenjun Zeng, Yan Lu

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(东部技术研究所) Microsoft Research Asia(微软亚洲研究院)

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20119 2025-09-25 cs.CV 70%

A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA

Belal Shoer, Yova Kementchedjhieva

机构 * MBZUAI(穆罕默德·本·拉希德智能研究院)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted at WiNLP, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16673 2025-09-23 cs.CV 70%

MedCutMix: A Data-Centric Approach to Improve Radiology Vision-Language Pre-training with Disease Awareness

Sinuo Wang, Yutong Xie, Yuyuan Liu, Qi Wu

机构 * The University of Adelaide(阿德莱德大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) University of Oxford(牛津大学)

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16560 2025-09-23 cs.CV 70%

Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

Ji Soo Lee, Byungoh Ko, Jaewon Cho, Howoong Lee, Jaewoon Byun, Hyunwoo J. Kim

机构 * Korea University(韩国大学) Hanwha Vision(翰威英航) KAIST(韩国科学技术院)

专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏