arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-20 至 2025-11-20 共收录 60 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 5 篇

2511.13889 2025-11-20 cs.CV cs.LG 70%

Uni-Hema: Unified Model for Digital Hematopathology

Abdul Rehman, Iqra Rasool, Ayisha Imran, Mohsen Ali, Waqas Sultani

机构 * Information Technology University of Punjab(旁遮普信息科技大学) Chughtai Lab(楚格塔实验室)

专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24473 2025-11-20 cs.CV cs.AI cs.CL cs.LG 67%

Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks

Shijie Lian, Changti Wu, Laurence Tianruo Yang, Hang Yuan, Bin Yu, Lei Zhang, Kai Chen

机构 * Huazhong University of Science and Technology(华中科技大学) Zhongguancun Academy(中关村学院) East China Normal University(华东师范大学) Zhengzhou University(郑州大学) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15633 2025-11-20 cs.CV cs.LG 57%

Hierarchical Semantic Tree Anchoring for CLIP-Based Class-Incremental Learning

基于CLIP的层次语义树锚定的类增量学习

Tao Hu, Lan Li, Zhen-Hao Xie, Da-Wei Zhou

机构 * School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学) State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

AI总结 HASTEN通过层次语义树锚定方法,有效解决CLIP类增量学习中的灾难性遗忘问题,提升模型对层次化类别结构的保持能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15515 2025-11-20 cs.CV 57%

Multi-Text Guided Few-Shot Semantic Segmentation

多文本引导的少样本语义分割

Qiang Jiao, Bin Yan, Yi Yang, Mengrui Shi, Qiang Zhang

机构 * State Key Laboratory of Electromechanical Integrated Manufacturing of High-Performance Electronic Equipments(高性能电子设备机电一体化制造国家重点实验室) Center for Complex Systems(复杂系统中心) School of Mechano-Electronic Engineering(机械电子工程学院) Xidian University(西安电子科技大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

AI总结 MTGNet通过融合多文本提示提升少样本语义分割性能,采用多文本先验细化、文本锚点特征融合和前景置信度加权注意力模块,有效增强分割鲁棒性与语义一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14901 2025-11-20 cs.CV 57%

FarSLIP: Discovering Effective CLIP Adaptation for Fine-Grained Remote Sensing Understanding

Zhenshi Li, Weikang Yu, Dilxat Muhtar, Xueliang Zhang, Pengfeng Xiao, Pedram Ghamisi, Xiao Xiang Zhu

机构 * Nanjing University(南京大学) Technical University of Munich(慕尼黑技术大学) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍兹中心)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 2 篇

2511.15312 2025-11-20 cs.CV 79%

A Multimodal Transformer Approach for UAV Detection and Aerial Object Recognition Using Radar, Audio, and Video Data

Mauro Larrat, Claudomiro Sales

机构 * Institute of Exact and Natural Sciences(精确与自然科学研究所) Federal University of Pará(巴西亚马逊联邦大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14764 2025-11-20 cs.IR cs.AI 57%

Image-Seeking Intent Prediction for Cross-Device Product Search

Mariya Hendriksen, Svitlana Vakulenko, Jordan Massiah, Gabriella Kazai, Emine Yilmaz

机构 * University of Oxford(牛津大学) Vienna University of Economics and Business(维也纳经济与商业大学) Amazon(亚马逊公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Oral at RecSys Gen AI for E-commerce 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2404.03179 2025-11-20 cs.CV cs.MM cs.SD eess.AS 82%

UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization

Tiantian Geng, Teng Wang, Jinming Duan, Yanfu Zhang, Weili Guan, Feng Zheng, Ling shao

机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学) School of Computer Science, University of Birmingham(计算机科学学院,伯明翰大学) Department of Computer Science, University of Hong Kong(计算机科学系,香港大学) Division of Informatics, Imaging and Data Sciences, University of Manchester(信息学、成像与数据科学系,曼彻斯特大学) William and Mary(威廉与玛丽学院) Harbin Institute of Technology(哈尔滨工业大学) UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学-Terminus AI实验室)

专题命中 视频多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Published on IEEE TPAMI

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 10280-10294, August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14770 2025-11-20 cs.IR cs.AI 79%

ExplainRec: Towards Explainable Multi-Modal Zero-Shot Recommendation with Preference Attribution and Large Language Models

Bo Ma, LuYao Liu, ZeHua Hu, Simon Lau

机构 * Department of Software \& Microelectronics Peking University Beijing, China Department of Software \& Microelectronics Peking University Beijing, China hangli\ Department of Software \& Microelectronics Peking University Beijing, China zehua\ Department of Software \& Microelectronics Peking University Beijing, China xiaofan\ Economic Law School China University of Political Science School of Computer Science Peking University Beijing, China

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10210 2025-11-20 cs.CV 79%

MK-SGN: A Spiking Graph Convolutional Network with Multimodal Fusion and Knowledge Distillation for Skeleton-based Action Recognition

Naichuan Zheng, Hailun Xia, Zeyu Liang, Yuchen Du

机构 * Beijing Laboratory of Advanced Information Networks, Beijing Key Laboratory of Network System Architecture and Convergence, School of Information and Communication Engineering, Beijing University of Posts and Telecommunications(北京先进信息网络实验室、网络系统架构与收敛重点实验室、信息与通信工程学院、北京邮电大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15342 2025-11-20 cs.HC cs.AI 74%

Reflexive Evidence-Based Multimodal Learning for Clean Energy Transitions: Causal Insights on Cooking Fuel Access, Urbanization, and Carbon Emissions

Shan Shan

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24709 2025-11-20 cs.CV 57%

IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?

Yang Chen, Minghao Liu, Yufan Shen, Yunwen Li, Tianyuan Huang, Xinyu Fang, Tianyu Zheng, Wenxuan Huang, Cheng Yang, Daocheng Fu, Jianbiao Mei, Rong Wu, Yunfei Zhao, Licheng Wen, Xuemeng Yang, Song Mao, Qunshu Lin, Zhi Yu, Yongliang Shen, Yu Qiao, Botian Shi

机构 * IWR-Bench Team(IWR-Bench团队)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 6 篇

2511.15435 2025-11-20 cs.CV cs.AI cs.IR 84%

HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation

HV-Attack:多模态检索增强生成的分层视觉攻击

Linyin Luo, Yujuan Ding, Yunshan Ma, Wenqi Fan, Hanjiang Lai

机构 * The Hong Kong Polytechnic University(香港理工大学) Sun Yat-Sen University(中山大学) Singapore Management University(新加坡管理学院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种分层视觉攻击方法,通过在图像输入中添加不可察觉扰动,破坏多模态检索增强生成系统的检索和生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19579 2025-11-20 cs.CV cs.AI cs.LG 81%

Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration

Francisco Mena, Dino Ienco, Cassio F. Dantas, Roberto Interdonato, Andreas Dengel

机构 * Department of Computer Science, University of Kaiserslautern-Landau (RPTU)(科斯拉尔特伦大学计算机科学系) SDS, German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) INRAE, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) CIRAD, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) INRIA, EVERGREEN, University of Montpellier(蒙彼利埃大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at the Machine Learning journal, CfP: Discovery Science 2024

Journal ref Machine Learning 114, 279 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15308 2025-11-20 cs.CV 70%

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques, Daniel Cremers

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Technical University of Munich(慕尼黑技术大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15488 2025-11-20 cs.CV cs.AI 62%

FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection

Jiangyong Yu, Changyong Shu, Sifan Zhou, Zichen Yu, Xing Hu, Yan Chen, Dawei Yang

机构 * HOUMO AI

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments This paper is acceptted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08008 2025-11-20 cs.AI 57%

Combining LLM Semantic Reasoning with GNN Structural Modeling for Multi-View Multi-Label Feature Selection

Zhiqi Chen, Yuzhou Liu, Jiarui Liu, Wanfu Gao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14160 2025-11-20 cs.CL cs.LG 57%

Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * Queen Mary University of London(伦敦大学玛丽女王学院) Institut Jožef Stefan(乔泽夫·斯蒂芬研究所)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2511.15030 2025-11-20 eess.SP 67%

WiCo-PG: Wireless Channel Foundation Model for Pathloss Map Generation via Synesthesia of Machines

Mingran Sun, Lu Bai, Ziwei Huang, Xuesong Cai, Xiang Cheng, Jianjun Wu

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15520 2025-11-20 cs.RO cs.AI 57%

Theoretical Closed-loop Stability Bounds for Dynamical System Coupled with Diffusion Policies

动态系统与扩散策略耦合下的理论闭环稳定性边界

Gabriel Lauzier, Alexandre Girard, François Ferland

机构 * Department of Mechanical Engineering, Universite de Sherbrooke(机械工程系, Sherbrooke 大学) Department of Electronic and Computer Engineering, Universite de Sherbrooke(电子与计算机工程系, Sherbrooke 大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文研究了动态系统与扩散策略耦合下的闭环稳定性边界,提出了一种更快的模仿学习框架和基于演示方差的稳定性判断指标。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02249 2025-11-20 cs.RO cs.AI 57%

Natural Selection via Foundation Models for Soft Robot Evolution

Changhe Chen, Xiaohao Xu, Xiangdong Wang, Xiaonan Huang

机构 * University of Michigan-Ann Arbor(密歇根大学安娜堡分校)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19646 2025-11-20 cs.LG 50%

Energy-based generator matching: A neural sampler for general state space

Dongyeop Woo, Minsu Kim, Minkyu Kim, Kiyoung Seong, Sungsoo Ahn

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15206 2025-11-20 cs.CR cs.IT math.IT 50%

Trustworthy GenAI over 6G: Integrated Applications and Security Frameworks

Bui Duc Son, Trinh Van Chien, Dong In Kim

专题命中 多模态生成 :multimodal(abstract)

Comments 8 pages, 5 figures. Submitted for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15026 2025-11-20 eess.SP 50%

WiCo-MG: Wireless Channel Foundation Model for Multipath Generation via Synesthesia of Machines

Zengrui Han, Lu Bai, Xuesong Cai, Xiang Cheng

专题命中 多模态生成 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14516 2025-11-20 cs.LG 50%

Full-Atom Peptide Design via Riemannian-Euclidean Bayesian Flow Networks

Hao Qian, Shikui Tu, Lei Xu

专题命中 多模态生成 :multimodal(abstract)

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 10 篇

2511.14439 2025-11-20 cs.CL 83%

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian, Jiayuan Chen, Wenrao Pang, Ruiyao Chen, Xinwei Peng, Renjie Lu, Sijie Ren, Guanxu Zhu, Xiaoqin Wu, Zhiqiang Liu, Rongzhao Zhang, Luyi Jiang, Bing Han, Yunqiu Wang, Jie Xu

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15059 2025-11-20 cs.CV cs.CL 81%

Evaluating Multimodal Large Language Models on Vertically Written Japanese Text

Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara

机构 * Waseda University(早稻田大学) NII(日本国立信息机构) NII LLMC Tokyo, Japan(日本国立信息机构LLMC东京)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 17pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18015 2025-11-20 cs.CV cs.AI 81%

Beyond Diagnosis: Evaluating Multimodal LLMs for Pathology Localization in Chest Radiographs

Advait Gosai, Arun Kavishwar, Stephanie L. McNamara, Soujanya Samineni, Renato Umeton, Alexander Chowdhury, William Lotter

机构 * University of California(加州大学) Dana-Farber Cancer Institute(达纳-法伯癌症研究所) Massachusetts General Hospital(麻省总医院) St. Jude Children’s Research Hospital(圣犹大儿童研究医院) Brigham and Women’s Hospital & Harvard Medical School(布里法伦医院及哈佛医学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Proceedings of the 5th Machine Learning for Health (ML4H) Symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15085 2025-11-20 cs.CV 79%

TiCAL:Typicality-Based Consistency-Aware Learning for Multimodal Emotion Recognition

Wen Yin, Siyu Zhan, Cencen Liu, Xin Hu, Guiduo Duan, Xiurui Xie, Yuan-Fang Li, Tao He

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07573 2025-11-20 cs.IR cs.CV 79%

A Hybrid Multimodal Deep Learning Framework for Intelligent Fashion Recommendation

Kamand Kalashi, Babak Teimourpour

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏