arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 105 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 12 篇

2410.17494 2025-11-11 eess.IV cs.CV 79%

Enhancing Multimodal Medical Image Classification using Cross-Graph Modal Contrastive Learning

Jun-En Ding, Chien-Chin Hsu, Chi-Hsiang Chu, Shuqiang Wang, Feng Liu

机构 * Department of Systems Engineering, Stevens Institute of Technology, Hoboken, New Jersey, USA(系统工程系,史蒂文斯理工学院) Department of Nuclear Medicine, Kaohsiung Chang Gung Memorial Hospital, Kaohsiung, Taiwan(高雄长庚纪念医院核医学部) Institute of Statistics, National University of Kaohsiung, Kaohsiung, Taiwan(国立高雄大学统计研究所) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China(深圳先进技术研究院,中国科学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03070 2025-11-11 cs.LG cs.MM 79%

FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization

Manh Duong Nguyen, Trung Thanh Nguyen, Huy Hieu Pham, Trong Nghia Hoang, Phi Le Nguyen, Thanh Trung Huynh

机构 * Hanoi University of Science and Technology(河内科学技术大学) Nagoya University(名古屋大学) Washington State University(华盛顿州立大学) Swiss Federal Institute of Technology Lausanne(洛桑联邦理工学院)

专题命中 跨模态检索 :cross-modal(title);multi-modal(abstract);分类 cs.MM

Comments The 22nd International Symposium on Network Computing and Applications (NCA 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06268 2025-11-11 cs.CV cs.CY 79%

LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval

Jian Zhang, Junyi Guo, Junyi Yuan, Huanda Lu, Yanlin Zhou, Fangyu Wu, Qiufeng Wang, Dongming Lu

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) NingboTech University(宁波科技学院) Dunhuang Academy(敦煌研究院) Zhejiang University(浙江大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05667 2025-11-11 cs.IR 78%

SARCH: Multimodal Search for Archaeological Archives

Nivedita Sinha, Bharati Khanijo, Sanskar Singh, Priyansh Mahant, Ashutosh Roy, Saubhagya Singh Bhadouria, Arpan Jain, Maya Ramanath

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06269 2025-11-11 cs.LG q-bio.QM 78%

LLM$^3$-DTI: A Large Language Model and Multi-modal data co-powered framework for Drug-Target Interaction prediction

Yuhao Zhang, Qinghong Guo, Qixian Chen, Liuwei Zhang, Hongyan Cui, Xiyi Chen

机构 * Polytechnic Institute(多技术学院) School of Pharmaceutical Sciences(药学学院) Innovation Center of Yangtze River Delta(长江三角洲创新中心) Agricultural Genomics Institute at Shenzhen(深圳农业基因组研究院) School of Public Health(公共卫生学院)

专题命中 跨模态检索 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05630 2025-11-11 q-bio.NC cs.AI 70%

BrainCSD: A Hierarchical Consistency-Driven MoE Foundation Model for Unified Connectome Synthesis and Multitask Brain Trait Prediction

Xiongri Shen, Jiaqi Wang, Yi Zhong, Zhenxi Song, Leilei Zhao, Liling Li, Yichen Wei, Lingyan Liang, Shuqiang Wang, Baiying Lei, Demao Deng, Zhiguo Zhang

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术系) School of Intelligence Science and Engineering, College of Artificial Intelligence, Harbin Institute of Technology(哈尔滨工业大学智能科学与工程学院) School of Biomedical Engineering, National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, Guangdong Key Laboratory for Biomedical, Measurements and Ultrasound Imaging, Shenzhen University Medical School, Shenzhen University(深圳大学医学院生物医学工程学院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20354 2025-11-11 cs.CL cs.AI 62%

Rethinking Text-based Protein Understanding: Retrieval or LLM?

Juntong Wu, Zijing Liu, He Cao, Hao Li, Bin Feng, Zishan Shu, Ke Yu, Li Yuan, Yu Li

机构 * Peking University, Shenzhen Graduate School(北京大学深圳研究生院) International Digital Economy Academy (IDEA)(国际数字经济学院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by Empirical Methods in Natural Language Processing 2025 (EMNLP 2025) Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06259 2025-11-11 cs.LG cs.AI 57%

Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra

Yiwen Zhang, Keyan Ding, Yihang Wu, Xiang Zhuang, Yi Yang, Qiang Zhang, Huajun Chen

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05894 2025-11-11 cs.CV 57%

Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning

Fei Yu, Quan Deng, Shengeng Tang, Yuehua Li, Lechao Cheng

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02073 2025-11-11 cs.AI 57%

Large model retrieval enhancement framework for construction site risk identification

Jiawei Li, Chengye Yang, Yaochen Zhang, Weilin Sun, Lei Meng, Xiangxu Meng

专题命中 跨模态检索 :image-text(abstract);分类 cs.AI

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态生成 13 篇

2511.06793 2025-11-11 cs.LG cs.AI 86%

Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models

Kunhao Li, Wenhao Li, Di Wu, Lei Yang, Jun Bai, Ju Jia, Jason Xue

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(title);分类 cs.AI

Comments Accepted at AAAI 2026 as a Conference Paper (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12658 2025-11-11 cs.DC 85%

HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving

Xianzhe Dong, Tongxuan Liu, Yuting Zeng, Liangyu Liu, Yang Liu, Siyu Wu, Yu Wu, Hailong Yang, Ke Zhang, Jing Li

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06284 2025-11-11 cs.CV cs.CL cs.MM 82%

Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective

Bing Wang, Ximing Li, Yanjun Wang, Changchun Li, Lin Yuanbo Wu, Buyu Wang, Shengsheng Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by AAAI 2026. 13 pages, 6 figures. Code: https://github.com/wangbing1416/RETSIMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10566 2025-11-11 cs.CV 79%

EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing

Umar Khalid, Kashif Munir, Hasan Iqbal, Azib Farooq, Jing Hua, Nazanin Rahnavard, Chen Chen, Victor Zhu, Zhengping Ji

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00596 2025-11-11 cs.CV 70%

Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control

Danfeng Li, Hui Zhang, Sheng Wang, Jiacheng Li, Zuxuan Wu

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) HiThink Research(HiThink研究机构)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22812 2025-11-11 cs.CL 57%

EditGRPO: Reinforcement Learning with Post-Rollout Edits for Clinically Accurate Chest X-Ray Report Generation

Kai Zhang, Christopher Malon, Lichao Sun, Martin Renqiang Min

机构 * NEC Laboratories America(NEC美国实验室) Lehigh University(莱特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18452 2025-11-11 cs.SD eess.AS 57%

DIFFA: Large Language Diffusion Models Can Listen and Understand

Jiaming Zhou, Hongjie Chen, Shiwan Zhao, Jian Kang, Jie Li, Enzhi Wang, Yujie Guo, Haoqin Sun, Hui Wang, Aobo Kong, Yong Qin, Xuelong Li

机构 * College of Computer Science, Nankai University(南开大学计算机科学学院) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究所)

专题命中 多模态生成 :multimodal(abstract);分类 eess.AS

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06422 2025-11-11 cs.CV 57%

DiffusionUavLoc: Visually Prompted Diffusion for Cross-View UAV Localization

Tao Liu, Kan Ren, Qian Chen

机构 * School of Electronic and Optical Engineering, Nanjing University of Science and Technology(电子与光学工程学院,南京理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15387 2025-11-11 cs.CV 57%

DIO: Refining Mutual Information and Causal Chain to Enhance Machine Abstract Reasoning Ability

Ruizhuo Song, Beiming Yuan

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments 15 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01322 2025-11-11 cs.CV 57%

FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors

Chenxi Li, Weijie Wang, Qiang Li, Bruno Lepri, Nicu Sebe, Weizhi Nie

机构 * Tianjin University(天津大学) University of Trento(特伦托大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

Comments Accepted by ACMMM2025, Our project webpage: https://tjulcx.github.io/FreeInsert/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05606 2025-11-11 cs.CV 57%

FreeBlend: Advancing Concept Blending with Staged Feedback-Driven Interpolation Diffusion

Yufan Zhou, Haoyu Shen, Huan Wang

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of Science and Technology of China(中国科学技术大学) Westlake University(西湖大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Webpage: https://petershen-csworld.github.io/FreeBlend

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06892 2025-11-11 cs.RO 50%

Multi-Agent AI Framework for Road Situation Detection and C-ITS Message Generation

Kailin Tong, Selim Solmaz, Kenan Mujkic, Gottfried Allmer, Bo Leng

机构 * Virtual Vehicle Research GmbH(虚拟车辆研究有限公司) ASFINAG Maut Service GmbH(ASFINAG收费服务有限公司) Tongji University(同济大学)

专题命中 多模态生成 :multimodal(abstract)

Comments submitted to TRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10684 2025-11-11 cs.LG math.OC stat.CO stat.ML 50%

MDNS: Masked Diffusion Neural Sampler via Stochastic Optimal Control

Yuchen Zhu, Wei Guo, Jaemoo Choi, Guan-Horng Liu, Yongxin Chen, Molei Tao

机构 * Georgia Institute of Technology(佐治亚理工学院) FAIR at Meta(Meta的FAIR)

专题命中 多模态生成 :multi-modal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态评测 20 篇

2511.06722 2025-11-11 cs.CV cs.AI cs.CL 85%

Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View

Jianyu Qi, Ding Zou, Wenrui Yan, Rui Ma, Jiaxu Li, Zhijie Zheng, Zhiguo Yang, Rongchang Zhao

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accpeted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19875 2025-11-11 cs.CV 83%

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

Kirolos Ataallah, Eslam Abdelrahman, Mahmoud Ahmed, Chenhui Gou, Khushbu Pahwa, Jian Ding, Mohamed Elhoseiny

机构 * KAUST(卡塔尔科技大学) Monash University(墨尔本大学) RICE University(里士满大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted for oral presentation at the EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07010 2025-11-11 cs.CL cs.CV cs.HC 81%

A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation

Siddharth Betala, Kushan Raj, Vipul Betala, Rohan Saswade

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at The 12th Workshop on Asian Translation, co-located with IJCLNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20665 2025-11-11 cs.CV cs.MM 81%

SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection

Yuxuan Li, Xiang Li, Yunheng Li, Yicheng Zhang, Yimian Dai, Qibin Hou, Ming-Ming Cheng, Jian Yang

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted as Oral in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05883 2025-11-11 cs.AI 79%

Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks

Hehai Lin, Hui Liu, Shilei Cao, Jing Li, Haoliang Li, Wenya Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) City University of Hong Kong(香港城市大学) Sun Yat-sen University(中山大学) Harbin Institute of Technology(哈尔滨工业大学) Nanyang Technological University(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18004 2025-11-11 cs.CV 79%

SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning

Yuhao Shen, Liyuan Sun, Yan Xu, Wenbin Liu, Shuping Zhang, Shawn Afvari, Zhongyi Han, Jiaoyan Song, Yongzhi Ji, Tao Lu, Xiaonan He, Xin Gao, Juexiao Zhou

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK–Shenzhen)(数据科学学院,香港中文大学(深圳)) Computer Science Program, CEMSE Division, King Abdullah University of Science and Technology (KAUST)(计算机科学项目,科学与工程学院,国王 Abdullah 科学技术大学) Center of Excellence on Smart Health, KAUST(智能健康卓越中心,国王 Abdullah 科学技术大学) Center of Excellence for Generative AI, KAUST(生成式人工智能卓越中心,国王 Abdullah 科学技术大学) Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University(皮肤科,北京安贞医院,首都医科大学) Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital(皮肤科,天津整合皮肤科研究院,天津中医药大学附属医院) Department of Dermatology, Beijing Aerospace General Hospital(皮肤科,北京航天总医院) Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院) DermAssure, LLC(DermAssure 公司) School of Medicine, New York Medical College(医学院,纽约医学院) Capital Medical University(首都医科大学) Department of Dermatology, Second Hospital of Jilin University(皮肤科,吉林大学第二医院) Emergency Critical Care Center, Beijing AnZhen Hospital, Capital Medical University(急诊重症中心,北京安贞医院,首都医科大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07276 2025-11-11 cs.LG 78%

RobustA: Robust Anomaly Detection in Multimodal Data

Salem AlMarri, Muhammad Irzam Liaqat, Muhammad Zaigham Zaheer, Shah Nawaz, Karthik Nandakumar, Markus Schedl

机构 * Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) IMT School for Advanced Studies(IMT高级研究学院) Johannes Kepler University Linz(林茨约瑟夫·冯·克莱门茨大学) Human-centered AI Group, AI Lab, Linz Institute of Technology(以人为中心的人工智能小组、人工智能实验室、林茨技术研究所)

专题命中 多模态评测 :multimodal(title,abstract)

Comments Submitted to IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏