arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-21 至 2025-10-21 共收录 94 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 13 篇

2507.09966 2025-10-21 eess.IV cs.AI cs.CV cs.LG 84%

Multimodal Fusion at Three Tiers: Physics-Driven Data Generation and Vision-Language Guidance for Brain Tumor Segmentation

Mingda Zhang

机构 * Software School, Yunnan University, Kunming 650504, Yunnan, China(云南大学软件学院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 31 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16198 2025-10-21 cs.CL 79%

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad, Hesham Abdelgawad, Mohamed Aref, Ahmed Fares

机构 * Department of Electrical Engineering, Faculty of Engineering at Shoubra, Benha University, Cairo 11629, Egypt(电气工程系,谢布拉工程学院,本海大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16036 2025-10-21 cs.CV 79%

IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection

Zewen Li, Zitong Yu, Qilang Ye, Weicheng Xie, Wei Zhuo, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院) College of Computer Science, Nankai University(南开大学计算机学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室) National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Instrumentation and Measurement (TIM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16870 2025-10-21 cs.CV 70%

Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding

Yudan Ren, Xinlong Wang, Kexin Wang, Tian Xia, Zihan Ma, Zhaowei Li, Xiangrong Bi, Xiao Li, Xiaowei He

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13946 2025-10-21 cs.AI 70%

Visual Instruction Bottleneck Tuning

Changdae Oh, Jiatong Li, Shawn Im, Sharon Li

机构 * Department of Computer Sciences, University of Wisconsin–Madison(计算机科学系,威斯康星大学麦迪逊分校)

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17771 2025-10-21 cs.AI cs.CV 62%

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo, Xianfeng Tang, Suhang Wang, Joy Zeng, Zhenwei Dai, Zhan Shi, Tianxin Wei, Benoit Dumoulin, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊) Penn State University(宾夕法尼亚州立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 10 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17651 2025-10-21 cs.CV cs.AI cs.LG 62%

Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs

Sébastien Thuau, Siba Haidar, Ayush Bajracharya, Rachid Chelouah

机构 * esieaLab(esiea实验室) ESIEA(ESIEA学院) ETIS Laboratory(ETIS实验室) CNRS(法国国家科学研究中心) UMR8051(UMR8051研究中心) University of CY Cergy(CY塞克大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 7 pages, 1 figure, FLTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17405 2025-10-21 cs.CL cs.AI 62%

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

Mardiyyah Oduwole, Prince Mireku, Fatimo Adebanjo, Oluwatosin Olajide, Mahi Aminu Aliyu, Jekaterina Novikova

机构 * ML Collective Ashesi University(阿什esi大学) Abubakar Tafawa Balewa University(阿布巴克尔·塔法瓦·巴勒瓦大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16973 2025-10-21 cs.CV cs.AI physics.med-ph 62%

Foundation Models in Medical Image Analysis: A Systematic Review and Meta-Analysis

Praveenbalaji Rajendran, Mojtaba Safari, Wenfeng He, Mingzhe Hu, Shansong Wang, Jun Zhou, Xiaofeng Yang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15430 2025-10-21 cs.CV cs.AI 62%

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

Shuang Liang, Zhihao Xu, Jialing Tao, Hui Xue, Xiting Wang

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Withdrawn due to an accidental duplicate submission. This paper (arXiv:2510.15430) was unintentionally submitted as a new entry instead of a new version of our previous work (arXiv:2508.09201)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14807 2025-10-21 eess.IV cs.AI cs.CV 62%

FetalCLIP: A Visual-Language Foundation Model for Fetal Ultrasound Image Analysis

Fadillah Maani, Numan Saeed, Tausifa Saleem, Zaid Farooq, Hussain Alasmawi, Werner Diehl, Ameera Mohammad, Gareth Waring, Saudabi Valappi, Leanne Bricker, Mohammad Yaqub

机构 * Department of Computer Vision(计算机视觉系) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫比兹人工智能大学) Department of Machine Learning(机器学习系) Corniche Hospital, Abu Dhabi Health Services Company (SEHA)(阿布扎赫尔医院,阿布扎赫健康服务公司(SEHA))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17777 2025-10-21 cs.CV 57%

SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference

Samir Khaki, Junxian Guo, Jiaming Tang, Shang Yang, Yukang Chen, Konstantinos N. Plataniotis, Yao Lu, Song Han, Zhijian Liu

机构 * NVIDIA MIT(麻省理工学院) UC San Diego(南加州大学圣地亚哥分校) University of Toronto(多伦多大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16394 2025-10-21 eess.IV 50%

FSAR-Cap: A Fine-Grained Two-Stage Annotated Dataset for SAR Image Captioning

Jinqi Zhang, Lamei Zhang, Bin Zou

专题命中 图文多模态 :image-text(abstract)

Comments 5pages,4figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2510.16437 2025-10-21 eess.AS 79%

Audio-Visual Speech Enhancement for Spatial Audio - Spatial-VisualVoice and the MAVE Database

Danielle Yaffe, Ferdinand Campe, Prachi Sharma, Dorothea Kolossa, Boaz Rafaely

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16893 2025-10-21 cs.SD cs.AI cs.CL eess.AS 67%

Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations

Bo-Han Feng, Chien-Feng Liu, Yu-Hsuan Li Liang, Chih-Kai Yang, Szu-Wei Fu, Zhehuai Chen, Ke-Han Lu, Sung-Feng Huang, Chao-Han Huck Yang, Yu-Chiang Frank Wang, Yun-Nung Chen, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学) NVIDIA

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01879 2025-10-21 cs.MM cs.CV cs.SD eess.AS 67%

Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision

Che Liu, Yingji Zhang, Dong Zhang, Weijie Zhang, Chenggong Gong, Yu Lu, Shilin Zhou, Ziliang Gan, Ziao Wang, Haipang Wu, Ji Liu, André Freitas, Qifan Wang, Zenglin Xu, Rongjuncheng Zhang, Yong Dai

机构 * Imperial College London(伦敦帝国学院) University of Manchester(曼彻斯特大学) HiThink Research(HiThink研究院) Soochow University(苏州大学) Hong Kong Baptist University(香港 Baptist大学) Idiap Research Institute(Idiap研究 institute) Meta AI Fudan University(复旦大学)

专题命中 音频语音多模态 :omni-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Project: https://github.com/HiThink-Research/NEXUS-O

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17617 2025-10-21 cs.HC cs.CV 57%

ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input

Hendric Voss, Stefan Kopp

机构 * Social Cognitive Systems Group, Bielefeld University(比勒菲尔德大学社会认知系统小组)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 10 篇

2510.17023 2025-10-21 cs.CV cs.MM 84%

Enrich and Detect: Video Temporal Grounding with Multimodal LLMs

Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras

机构 * FAIR, Meta(FAIR、Meta) Johns Hopkins University(约翰霍普金斯大学) Northeastern University(东北大学)

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments ICCV 2025 (Highlights)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17038 2025-10-21 cs.RO cs.AI cs.CV 81%

DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation

Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi

机构 * Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学) The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院)) Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16972 2025-10-21 cs.CV cs.AI 73%

The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yikang Zhou, Haobo Yuan, Lu Qi, Xiangtai Li, Shunping Ji

机构 * Wuhan University(武汉大学) University of California, Merced(加州大学默塞德分校) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments The 1st place report of 7th LSVOS challenge RVOS track in ICCV 2025. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16444 2025-10-21 cs.CV cs.MM cs.RO eess.IV 62%

RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba

Kunyu Peng, Di Wen, Jia Fu, Jiamin Wu, Kailun Yang, Junwei Zheng, Ruiping Liu, Yufan Chen, Yuqian Fu, Danda Pani Paudel, Luc Van Gool, Rainer Stiefelhagen

机构 * Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology(人机化研究所,卡尔斯鲁厄技术大学) RISE Research Institutes of Sweden(瑞典RISE研究机构) KTH Royal Institute of Technology(皇家理工学院) School of Artificial Intelligence and Robotics(人工智能与机器人学院) National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) Chinese University of Hong Kong(香港中文大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Extended version of ECCV 2024 paper arXiv:2407.01872. The dataset and code are released at https://github.com/KPeng9510/refAVA2

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17384 2025-10-21 cs.CV 57%

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17212 2025-10-21 cs.LG cs.AI 57%

D2C-HRHR: Discrete Actions with Double Distributional Critics for High-Risk-High-Return Tasks

Jundong Zhang, Yuhui Situ, Fanji Zhang, Rongji Deng, Tianqi Wei

机构 * School of Artificial Intelligence, Sun Yat-sen University(人工智能学院,中山大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16989 2025-10-21 cs.CV 57%

Training-free Online Video Step Grounding

Luca Zanella, Massimiliano Mancini, Yiming Wang, Alessio Tonioni, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会) Google(谷歌)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025. Project website at https://lucazanella.github.io/baglm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16455 2025-10-21 cs.CL 57%

RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning

Deyi Ji, Yuekui Yang, Haiyang Wu, Shaoping Ma, Tianrun Chen, Lanyun Zhu

机构 * Tencent(腾讯公司) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Zhejiang University(浙江大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments ACL 2025 (Oral, Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16980 2025-10-21 cs.LG 50%

Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision

Kanghui Ning, Zijie Pan, Yushan Jiang, Anderson Schneider, Yuriy Nevmyvaka, Dongjin Song

机构 * School of Computing University of Connecticut Storrs, CT(计算学院 美国康涅狄格大学 斯托尔斯分校) Department of Machine Learning Research Morgan Stanley New York, NY(机器学习研究部 花旗集团 新 York)

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19949 2025-10-21 eess.IV cs.LG 50%

Automated Video-EEG Analysis in Epilepsy Studies: Advances and Challenges

Valerii A. Zuev, Elena G. Salmagambetova, Stepan N. Djakov, Lev V. Utkin

机构 * Peter the Great St.Petersburg Polytechnic University(彼得大帝圣彼得堡理工大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2510.14605 2025-10-21 cs.CV cs.AI 81%

Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Alibaba Cloud Computing(阿里巴巴云计算)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17415 2025-10-21 cs.CL cs.AI cs.MA cs.MM cs.SE 67%

BenCao: An Instruction-Tuned Large Language Model for Traditional Chinese Medicine

Jiacheng Xie, Yang Yu, Yibo Chen, Hanyao Zhang, Lening Zhao, Jiaxuan He, Lei Jiang, Xiaoting Tang, Guanghui An, Dong Xu

机构 * Community Health Service Center Shanghai Pudong New Area(上海浦东新区社区卫生服务中心) School of Acupuncture-Moxibustion and Tuina, Shanghai University of Traditional Chinese Medicine(上海中医药大学针灸推拿学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17685 2025-10-21 cs.CV cs.AI 62%

Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning

Min Cao, Xinyu Zhou, Ding Jiang, Bo Du, Mang Ye, Min Zhang

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) School of Computer Science, Wuhan University(武汉大学计算机学院) Key Laboratory of New Generation Artificial Intelligence Technology & Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新 generation 人工智能技术及交叉应用重点实验室(东南大学),教育部,中国)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Final version published in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Xplore link: https://ieeexplore.ieee.org/document/11199360

详情

展开后加载摘要…

URL PDF HTML 收藏