arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 15 篇

2511.07080 2025-11-11 cs.CL cs.AI 84%

Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora

Khalil Hennara, Ahmad Bastati, Muhammad Hreden, Mohamed Motasim Hamed, Zeina Aldallal, Sara Chrouf, Safwan AlModhayan

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06665 2025-11-11 cs.CV cs.AI 81%

Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks

Lingran Song, Yucheng Zhou, Jianbing Shen

机构 * Lingran Song, Yucheng Zhou, Jianbing Shen(作者)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05577 2025-11-11 cs.LG cond-mat.mtrl-sci cs.AI cs.CL 81%

Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction

An Vuong, Minh-Hao Van, Prateek Verma, Chen Zhao, Xintao Wu

机构 * Department of EECS University of Arkansas(电子工程与科学系 亚拉荷加大学) Department of CS Baylor University(计算机科学系 基尔默大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09997 2025-11-11 cs.CV 79%

Descriptive Image-Text Matching with Graded Contextual Similarity

Jinhyun Jang, Jiyoung Lee, Kwanghoon Sohn

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments This version is incomplete and requires substantial revisions and extensions. We withdraw the paper and plan to submit a thoroughly revised version as a new submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20013 2025-11-11 cs.LG cs.AI cs.IR 79%

Cross-Platform E-Commerce Product Categorization and Recategorization: A Multimodal Hierarchical Classification Approach

Lotte Gross, Rebecca Walter, Nicole Zoppi, Adrien Justus, Alessandro Gambetti, Qiwei Han, Maximilian Kaiser

机构 * Nova School of Business and Economics(诺瓦商业与经济学院) Nova School of Business(诺瓦商业学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accetped at IEEE BigData 2025, 10 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02113 2025-11-11 cs.IR 78%

Enhancing Multimodal Recommendations with Vision-Language Models and Information-Aware Fusion

Hai-Dang Kieu, Min Xu, Thanh Trung Huynh, Dung D. Le

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11196 2025-11-11 cs.CL cs.CV 76%

Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations

Johannes Moll, Markus Graf, Tristan Lemke, Nicolas Lenhart, Daniel Truhn, Jean-Benoit Delbrouck, Jiazhen Pan, Daniel Rueckert, Lisa C. Adams, Keno K. Bressem

机构 * Technical University of Munich (TUM)(慕尼黑技术大学) TUM University Hospital(慕尼黑技术大学医院) German Heart Center TUM University Hospital(慕尼黑技术大学医院德国心脏中心) Department of Radiology(放射科) Klinikum rechts der Isar TUM University Hospital(慕尼黑技术大学医院右岸诊所) Uniklinik RWTH Aachen(亚琛工业大学医院) HOPPR IL USA(HOPPR美国) University of Oxford(牛津大学) Imperial College London(伦敦帝国学院)

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL

Comments Accepted to ML4H 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02044 2025-11-11 cs.CV 74%

A multi-modal vision-language model for generalizable annotation-free pathology localization

Hao Yang, Hong-Yu Zhou, Jiarun Liu, Weijian Huang, Cheng Li, Zhihuan Li, Yuanxu Gao, Qiegen Liu, Yong Liang, Qi Yang, Song Wu, Tao Tan, Hairong Zheng, Kang Zhang, Shanshan Wang

机构 * Paul C. Lauterbur Research Center for Biomedical Imaging(生物医学成像研究中心) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Pengcheng Laboratory(鹏城实验室) University of Chinese Academy of Sciences(中国科学院大学) Chinese Medicine Guangdong Laboratory(广东中医药实验室) Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院)

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06947 2025-11-11 cs.CV cs.AI 73%

FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection

Yulin Chen, Zeyuan Wang, Tianyuan Yu, Yingmei Wei, Liang Bai

机构 * Laboratory for Big Data and Decision, National University of Defense Technology, Changsha 410073, China(大数据与决策实验室,国防科技大学,长沙410073,中国)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 15 page, 9 figures, published to PRCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06653 2025-11-11 cs.CV cs.CL 73%

HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment

Ruijia Wu, Ping Chen, Fei Shen, Shaoan Zhao, Qiang Hui, Huanlin Gao, Ting Lu, Zhaoxiang Liu, Fang Zhao, Kai Wang, Shiguo Lian

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted by AAAI 2026 as an Oral Presentation (13 pages, 7 figures, 7 tables)

Journal ref AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07171 2025-11-11 cs.CV cs.AI cs.LG 62%

Federated Learning for Video Violence Detection: Complementary Roles of Lightweight CNNs and Vision-Language Models for Energy-Efficient Use

Sébastien Thuau, Siba Haidar, Rachid Chelouah

机构 * esieaLab, ETIS Laboratory(esiea实验室,ETIS实验室) ESIEA, University of CY Cergy(ESIEA,CY塞克大学) ETIS Laboratory, CNR1S, UMR8051(ETIS实验室,CNR1S,UMR8051)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 3 figures, ICTAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06908 2025-11-11 cs.CV cs.MM 62%

Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding

Yuzhen Li, Min Liu, Zhaoyang Li, Yuan Bian, Xueping Wang, Erbo Zhai, Yaonan Wang

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15052 2025-11-11 cs.CL cs.AI 62%

Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning

Siddharth Betala, Ishan Chokshi

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted at the Ninth Conference on Machine Translation (WMT24), co-located with EMNLP 2024

Journal ref https://aclanthology.org/2024.wmt-1.81/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07068 2025-11-11 cs.CV cs.LG 57%

ClusterMine: Robust Label-Free Visual Out-Of-Distribution Detection via Concept Mining from Text Corpora

Nikolas Adaloglou, Diana Petrusheva, Mohamed Asker, Felix Michels, Markus Kollmann

机构 * Heinrich Heine University of Düsseldorf(海因里希-海涅大学杜塞尔多夫分校)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted in WACV 2026. Code in https://github.com/HHU-MMBS/clustermine_wacv_official 9 Tables, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05642 2025-11-11 cs.RO cs.AR cs.CV cs.SY eess.SY 57%

Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots

Justin Williams, Kishor Datta Gupta, Roy George, Mrinmoy Sarkar

机构 * Department of Cyber-Physical Systems, Clark Atlanta University(克雷克阿特拉大学计算机物理系统系) Siemens Corporation(西门子公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏