arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-09 至 2025-09-09 共收录 75 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 8 篇

2406.06620 2025-09-09 cs.LG cs.AI cs.CL 84%

MedualTime: A Dual-Adapter Language Model for Medical Time Series-Text Multimodal Learning

Jiexia Ye, Weiqi Zhang, Ziyue Li, Jia Li, Meng Zhao, Fugee Tsung

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Technical University of Munich(慕尼黑技术大学) Columbia University(哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 6 figure, 3 tables

Journal ref IJCAI 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06291 2025-09-09 cs.CV 83%

Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding

Jiangnan Xie, Xiaolong Zheng, Liang Zheng

机构 * College of Electronics and Information, Hangzhou Dianzi University(电子信息学院,杭州电子大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05615 2025-09-09 cs.LG cs.AI 79%

Causal Debiasing Medical Multimodal Representation Learning with Missing Modalities

Xiaoguang Zhu, Lianlong Sun, Yang Liu, Pengyi Jiang, Uma Srivatsa, Nipavan Chiamvimonvat, Vladimir Filkov

机构 * DataLab: Data Science and Informatics, University of California, Davis(加州大学戴维斯分校数据实验室) Department of Electrical and Computer Engineering, University of Rochester(罗切斯特大学电气与计算机工程系) Academy for Engineering & Technology, Fudan University(复旦大学工程与技术学院) Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Department of Electrical and Computer Engineering, New York University(纽约大学电气与计算机工程系) UC Davis Health(加州大学戴维斯分校医疗中心) Department of Basic Medical Sciences, University of Arizona(亚利桑那大学基础医学系) Department of Computer Science, University of California, Davis(加州大学戴维斯分校计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Submitted to IEEE TKDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04139 2025-09-09 cs.CV cs.AI cs.ET cs.LG cs.RO 73%

Driver-Net: Multi-Camera Fusion for Assessing Driver Take-Over Readiness in Automated Vehicles

Mahdi Rezaei, Mohsen Azarmi

机构 * Institute for Transport Studies, Computer Vision and Machine Learning Group, University of Leeds(交通研究 institute,计算机视觉和机器学习组,莱斯特大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Journal ref 2025 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06321 2025-09-09 cs.CV 70%

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling

Mengcheng Lan, Chaofeng Chen, Jiaxing Xu, Zongrui Li, Yiping Ke, Xudong Jiang, Yingchen Yu, Yunqing Zhao, Song Bai

机构 * College of Computing and Data Science, Nanyang Technological University(computing and Data Science学院,南洋理工大学) School of Electrical and Electronic Engineering, Nanyang Technological University(Electrical and Electronic Engineering学院,南洋理工大学) School of Artificial Intelligence, Wuhan University(Artificial Intelligence学院,武汉大学) ByteDance(字节跳动)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Extended version of our conference paper arXiv:2410.09855

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11299 2025-09-09 cs.CL cs.AI 62%

BriLLM: Brain-inspired Large Language Model

Hai Zhao, Hongqiu Wu, Dongjie Yang, Anni Zou, Jiale Hong

机构 * AGI Institute(AGI研究院) Computer School(计算机学院) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23067 2025-09-09 cs.AI 57%

FairReason: Balancing Reasoning and Social Bias in MLLMs

Zhenyu Pan, Yutong Zhang, Jianshu Zhang, Haoran Lu, Haozheng Luo, Yuwei Han, Philip S. Yu, Manling Li, Han Liu

机构 * Northwestern University(西北大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted to the Trustworthy FMs workshop in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15113 2025-09-09 cs.CY cs.AI cs.LG 57%

Transit for All: Mapping Equitable Bike2Subway Connection using Region Representation Learning

Min Namgung, JangHyeon Lee, Fangyi Ding, Yao-Yi Chiang

机构 * University of Minnesota, Twin Cities(明尼苏达大学,双城分校) The University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments SIGSPATIAL 25

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 7 篇

2504.00717 2025-09-09 cs.NE cs.AI 83%

Advancements in Multimodal Differential Evolution: A Comprehensive Review and Future Perspectives

Dikshit Chauhan, Shivani, Donghwi Jung, Anupam Yadav

机构 * National University of Singapore(新加坡国立大学) Dr. B.R. Ambedkar National Institute of Technology Jalandhar(德拉·B.R. 阿姆贝德卡国家理工学院贾兰德哈尔) Korea University(韩国大学)

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Journal ref Artificial Intelligence Review 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06312 2025-09-09 eess.SY cs.LG cs.SY 82%

Enhancing Low-Altitude Airspace Security: MLLM-Enabled UAV Intent Recognition

Guangyu Lei, Tianhao Liang, Yuqi Ping, Xinglin Chen, Longyu Zhou, Junwei Wu, Xiyuan Zhang, Huahao Ding, Xingjian Zhang, Weijie Yuan, Tingting Zhang, Qinyu Zhang

机构 * School of Information Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)信息科学与技术学院) Guangdong Provincial Key Laboratory of Space-Aerial Networking and Intelligent Sensing(广东省空间-航空网络与智能感知重点实验室) Information Systems Technology and Design, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计) School of System Design and Intelligent Manufacturing, Southern University of Science and Technology(南方科技大学系统设计与智能制造学院)

专题命中 其他多模态 :MLLM(title,abstract);multimodal(abstract)

Comments The paper has been submitted to IEEE Internet of Things Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06617 2025-09-09 eess.IV cs.CV 79%

MM-DINOv2: Adapting Foundation Models for Multi-Modal Medical Image Analysis

Daniel Scholz, Ayhan Can Erdur, Viktoria Ehm, Anke Meyer-Baese, Jan C. Peeken, Daniel Rueckert, Benedikt Wiestler

机构 * Chair for AI for Image-Guided Diagnosis and Therapy, Technical University of Munich (TUM)(人工智能辅助影像诊断与治疗研究所,慕尼黑技术大学) TUM University Hospital, Munich, Germany(慕尼黑技术大学医院) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM)(人工智能在医疗与健康领域研究所,慕尼黑技术大学) Department of Radiation Oncology, TUM University Hospital, Munich, Germany(放射肿瘤科,慕尼黑技术大学医院) Chair for Computer Vision and Artificial Intelligence, Technical University of Munich (TUM)(计算机视觉与人工智能研究所,慕尼黑技术大学) Department of Scientific Computing, Florida State University(科学计算系,佛罗里达州立大学) Deutsches Konsortium für Translationale Krebsforschung (DKTK), Partner Site Munich(德国转化癌症研究联盟(DKTK)慕尼黑分部) Institute of Radiation Medicine (IRM), Department of Radiation Sciences (DRS), Helmholtz Center Munich(放射医学研究所(IRM),辐射科学部门(DRS),海德堡中心慕尼黑) Institute for Advanced Study, Technical University of Munich (TUM)(高级研究所,慕尼黑技术大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06333 2025-09-09 cs.CV cs.RO 79%

Multi-Modal Camera-Based Detection of Vulnerable Road Users

Penelope Brown, Julie Stephany Berrio Perez, Mao Shan, Stewart Worrall

机构 * The University of Sydney(悉尼大学) Australian Centre for Robotics(澳大利亚机器人中心)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06219 2025-09-09 cs.LG cs.MM 79%

MCIGLE: Multimodal Exemplar-Free Class-Incremental Graph Learning

Haochen You, Baojing Liu

机构 * Graduate School of Arts and Sciences(艺术与科学研究生院) Columbia University, New York, USA(哥伦比亚大学) School of Artificial Intelligence(人工智能学院) Hebei Institute of Communications, Shijiazhuang, PR China(河北通信学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted as a conference paper at KSEM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05671 2025-09-09 cs.LG cs.AI cs.CR stat.ML 79%

GraMFedDHAR: Graph Based Multimodal Differentially Private Federated HAR

Labani Halder, Tanmay Sen, Sarbani Palit

机构 * Indian Statistical Institute Kolkata(印度统计研究所科塔加特)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06167 2025-09-09 cs.LG cs.GR 50%

Exploring Urban Factors with Autoencoders: Relationship Between Static and Dynamic Features

Ximena Pocco, Waqar Hassan, Karelia Salinas, Vladimir Molchanov, Luis G. Nonato

机构 * ICMC, University of Sao Paulo, Sao Carlos, Brazil(圣卡洛斯大学ICMC) Münster University, Westphalia, Germany(西发里亚大学)

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏