arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-01 至 2025-10-01 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2509.25717 2025-10-01 cs.CV cs.CL cs.LG 81%

Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization

Xintong Li, Chuhan Wang, Junda Wu, Rohan Surana, Tong Yu, Julian McAuley, Jingbo Shang

机构 * University of California, San Diego(加州大学圣迭戈分校) Adobe Research(Adobe研究)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22820 2025-10-01 cs.CV cs.AI 81%

MMPB: It's Time for Multi-Modal Personalization

Jaeik Kim, Woojin Kim, Woohyeon Park, Jaeyoung Do

机构 * AIDAS Laboratory(AIDAS实验室) IPAI ECE(电子工程系) Seoul National University(首尔国立大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25791 2025-10-01 cs.CV 79%

EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks

Yuan Gao, Sangwook Kim, Chris McIntosh

机构 * Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络) Department of Medical Biophysics, UofT(医学生物物理学系) Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络) Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学) Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute) Department of Medical Imaging, UofT(医学影像学系) Vector Institute, Toronto(向量研究所)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments MICCAI 2025

Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. MICCAI 2025. Lecture Notes in Computer Science, vol 15964. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19294 2025-10-01 cs.CV cs.AI cs.CL 78%

Object Detection with Multimodal Large Vision-Language Models: An In-depth Review

Ranjan Sapkota, Manoj Karkee

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments First Peer Reviewed Review Paper for Object Detection with Vision-Language Models (VLMs)

Journal ref Information Fusion, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25654 2025-10-01 cs.CV 70%

DescribeEarth: Describe Anything for Remote Sensing Images

Kaiyu Li, Zixuan Jiang, Xiangyong Cao, Jiayu Wang, Yuchen Xiao, Deyu Meng, Zhi Wang

机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室) School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室) Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)

专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06105 2025-10-01 cs.CV 70%

PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology

Yating Huang, Ziyan Huang, Lintao Xiang, Qijun Yang, Hujun Yin

机构 * University of Manchester(曼彻斯特大学) South China University of Technology(华南理工大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accept by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25202 2025-10-01 cs.LG 67%

VLHSA: Vision-Language Hierarchical Semantic Alignment for Jigsaw Puzzle Solving with Eroded Gaps

Zhuoning Xu, Xinyan Liu

机构 * Zhuoning Xu 1(Xu Zhuoning 1) Xinyan Liu 1(Liu Xinyan 1)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07675 2025-10-01 cs.LG cs.AI cs.CV 62%

Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization

Seongjae Kang, Dong Bok Lee, Hyungjoon Jang, Sung Ju Hwang

机构 * VUNO Inc.(VUNO公司) KAIST(韩国科学技术院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 38 pages, 17 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21432 2025-10-01 cs.RO cs.CV 57%

UAV-VLN: End-to-End Vision Language guided Navigation for UAVs

Pranav Saxena, Nishant Raghuvanshi, Neena Goveas

机构 * Birla Institute of Technology and Science Pilani, K.K Birla Goa Campus(比拉理工学院和科学学院,K.K比拉果阿校区)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Journal ref Proc. European Conference on Mobile Robots (ECMR), 2025, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏