arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-20 至 2025-08-20 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 11 篇

2508.13843 2025-08-20 cs.IR cs.AI 88%

UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion

Zihan Liang, Yufei Ma, ZhiPeng Qian, Huangyu Dai, Zihan Wang, Ben Chen, Chenyi Lei, Yuqing Ding, Han Li

机构 * Kuaishou Technology(快手科技)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

Comments Accepted at CIKM2025 as a long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13387 2025-08-20 cs.AI 85%

SPANER: Shared Prompt Aligner for Multimodal Semantic Representation

Thye Shan Ng, Caren Soyeon Han, Eun-Jung Holden

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13460 2025-08-20 cs.CV 83%

Revisiting MLLM Token Technology through the Lens of Classical Visual Coding

Jinming Liu, Junyan Lin, Yuntao Wei, Kele Shao, Keda Tao, Jianguo Huang, Xudong Yang, Zhibo Chen, Huan Wang, Xin Jin

机构 * Eastern Institute of Technology, Ningbo, China(东部技术研究所) Westlake University(西湖大学) USTC(中国科学技术大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13971 2025-08-20 eess.AS cs.CL cs.HC cs.LG cs.MM 82%

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience

Andrew Chang, Chenkai Hu, Ji Qi, Zhuojian Wei, Kexin Zhang, Viswadruth Akkaraju, David Poeppel, Dustin Freeman

机构 * New York UniversityUSA(纽约大学) Max Planck SocietyGermany(马克斯·普朗克研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.MM、eess.AS

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12992 2025-08-20 cs.MM 79%

MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex Scenarios

Xiangxian Li, Yawen Zheng, Baiqiao Zhang, Yijia Ma, Xianhui Cao, Juan Liu, Yulong Bian, Jin Huang, Chenglei Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13196 2025-08-20 cs.LG cs.AI cs.IR 79%

Contextual Attention-Based Multimodal Fusion of LLM and CNN for Sentiment Analysis

Meriem Zerkouk, Miloud Mihoubi, Belkacem Chikhaoui

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments The 38th Canadian Conference on Artificial Intelligence ( 2025 )

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16873 2025-08-20 cs.CV 74%

ContrastAlign: Toward Robust BEV Feature Alignment via Contrastive Learning for Multi-Modal 3D Object Detection

Ziying Song, Hongyu Pan, Feiyang Jia, Yongchang Zhang, Lin Liu, Lei Yang, Shaoqing Xu, Peiliang Wu, Caiyan Jia, Zheng Zhang, Yadan Luo

机构 * School of Computer Science & Technology, Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing Jiaotong University(计算机科学与技术学院,北京交通大数据挖掘与具身智能重点实验室,北京交通大学) Horizon Robotics Nanyang Technological University(南洋理工大学) University of Macau(澳门大学) School of Information Science and Engineering, Yanshan University(信息科学与工程学院,燕山大学) School of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术学院,哈尔滨理工大学) School of Information Technology and Electrical Engineering, The University of Queensland(信息技术与电气工程学院,昆士兰大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04107 2025-08-20 cs.CV cs.AI 73%

Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder

Jingchao Wang, Zhijian Wu, Dingjiang Huang, Yefeng Zheng, Hong Wang

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07796 2025-08-20 cs.CV cs.AI cs.LG 62%

Fusing Echocardiography Images and Medical Records for Continuous Patient Stratification

Nathan Painchaud, Jérémie Stym-Popper, Pierre-Yves Courand, Nicolas Thome, Pierre-Marc Jodoin, Nicolas Duchateau, Olivier Bernard

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages + 2 pages of supplementary material, accepted for publication in IEEE TUFFC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14015 2025-08-20 cs.CV 57%

Backdooring Self-Supervised Contrastive Learning by Noisy Alignment

Tuo Chen, Jie Gui, Minjing Dong, Ju Jia, Lanting Fang, Jian Liu

机构 * Southeast University(东南大学) Purple Mountain Laboratories(紫金山实验室) Ant Group(蚂蚁集团) City University of Hong Kong(香港城市大学) Beijing Institute of Technology(北京理工大学)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13995 2025-08-20 cs.CV 57%

Self-Supervised Sparse Sensor Fusion for Long Range Perception

Edoardo Palladin, Samuel Brucker, Filippo Ghilotti, Praveen Narayanan, Mario Bijelic, Felix Heide

机构 * Torc Robotics(Torc机器人公司) Princeton University(普林斯顿大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏