arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-20 至 2025-08-20 共收录 36 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 11 篇

2405.16873 2025-08-20 cs.CV 74%

ContrastAlign: Toward Robust BEV Feature Alignment via Contrastive Learning for Multi-Modal 3D Object Detection

Ziying Song, Hongyu Pan, Feiyang Jia, Yongchang Zhang, Lin Liu, Lei Yang, Shaoqing Xu, Peiliang Wu, Caiyan Jia, Zheng Zhang, Yadan Luo

机构 * School of Computer Science & Technology, Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing Jiaotong University(计算机科学与技术学院,北京交通大数据挖掘与具身智能重点实验室,北京交通大学) Horizon Robotics Nanyang Technological University(南洋理工大学) University of Macau(澳门大学) School of Information Science and Engineering, Yanshan University(信息科学与工程学院,燕山大学) School of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术学院,哈尔滨理工大学) School of Information Technology and Electrical Engineering, The University of Queensland(信息技术与电气工程学院,昆士兰大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04107 2025-08-20 cs.CV cs.AI 73%

Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder

Jingchao Wang, Zhijian Wu, Dingjiang Huang, Yefeng Zheng, Hong Wang

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07796 2025-08-20 cs.CV cs.AI cs.LG 62%

Fusing Echocardiography Images and Medical Records for Continuous Patient Stratification

Nathan Painchaud, Jérémie Stym-Popper, Pierre-Yves Courand, Nicolas Thome, Pierre-Marc Jodoin, Nicolas Duchateau, Olivier Bernard

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages + 2 pages of supplementary material, accepted for publication in IEEE TUFFC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14015 2025-08-20 cs.CV 57%

Backdooring Self-Supervised Contrastive Learning by Noisy Alignment

Tuo Chen, Jie Gui, Minjing Dong, Ju Jia, Lanting Fang, Jian Liu

机构 * Southeast University(东南大学) Purple Mountain Laboratories(紫金山实验室) Ant Group(蚂蚁集团) City University of Hong Kong(香港城市大学) Beijing Institute of Technology(北京理工大学)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13995 2025-08-20 cs.CV 57%

Self-Supervised Sparse Sensor Fusion for Long Range Perception

Edoardo Palladin, Samuel Brucker, Filippo Ghilotti, Praveen Narayanan, Mario Bijelic, Felix Heide

机构 * Torc Robotics(Torc机器人公司) Princeton University(普林斯顿大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 1 篇

2508.05068 2025-08-20 cs.CV cs.AI cs.LG eess.IV 62%

Automatic Image Colorization with Convolutional Neural Networks and Generative Adversarial Networks

Changyuan Qiu, Hangrui Cao, Qihan Ren, Ruiyu Li, Yuqing Qiu

机构 * University of Michigan(密歇根大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments All authors have equal authorship and equal contribution, ranked in alphabetic order. First version of this paper was completed and published in 2021

详情

展开后加载摘要…

URL PDF HTML 收藏