arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2603.11380 2026-03-26 cs.CV 83%

DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding

DriveXQA: 多模态视觉问答用于恶劣驾驶场景理解

Mingzhe Tao, Ruiping Liu, Junwei Zheng, Yufan Chen, Kedi Ying, M. Saquib Sarfraz, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Hunan University(湖南大学) Mercedes-Benz Tech Innovation(梅赛德斯-奔驰技术创新)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出DriveXQA多模态数据集,用于自主驾驶场景中的视觉问答,通过融合多传感器信息提升对异常驾驶场景的理解能力。

Comments Accepted to CVPR DriveX Workshop. Dataset and Code: https://github.com/jtjmd/DRIVEXQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23276 2026-03-25 cs.CV 83%

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

CCF:互补协作融合用于领域泛化的多模态3D目标检测

Yuchen Wu, Kun Wang, Yining Pan, Na Zhao

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出CCF方法,通过查询解耦损失、LiDAR引导深度先验和互补跨模态掩码,提升多模态3D目标检测在跨领域场景下的鲁棒性,实验表明优于现有方法且保持源域性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22852 2026-03-25 cs.CV 83%

Gau-Occ: Geometry-Completed Gaussians for Multi-Modal 3D Occupancy Prediction

Gau-Occ:用于多模态3D占用预测的几何完备高斯分布

Chengxin Lv, Yihui Li, Hongyu Yang, YunHong Wang

机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,北京航空航天大学,北京,中国) School of Computer Science and Engineering, Beihang University, Beijing, China(计算机科学与工程学院,北京航空航天大学,北京,中国) School of Artificial Intelligence, Beihang University, Beijing, China(人工智能学院,北京航空航天大学,北京,中国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 Gau-Occ通过几何完备的3D高斯分布实现多模态3D占用预测,利用LiDAR完成扩散器恢复缺失结构并融合多视角图像语义,提升空间一致性和语义区分性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21584 2026-03-24 cs.LG cs.CV 83%

SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models

SSAM:奇异子空间对齐用于融合多模态大语言模型

Md Kaykobad Reza, Ameya Patil, Edward Ayrapetian, M. Salman Asif

机构 * University of California Riverside(加州大学河滨分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SSAM通过参数空间对齐融合多模态大语言模型,无需训练数据实现跨模态统一,提升性能并降低资源消耗。

Comments 25 Pages, 9 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14188 2026-03-24 cs.CV 83%

Joint Segmentation and Grading with Iterative Optimization for Multimodal Glaucoma Diagnosis

多模态青光眼诊断的联合分割与分级迭代优化方法

Zhiwei Wang, Yuxing Li, Meilu Zhu, Defeng He, Edmund Y. Lam

机构 * Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong, China(香港大学电子与电气工程系) College of Information Engineering, Zhejiang University of Technology, Hangzhou, China(浙江工业大学信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种迭代多模态优化模型,通过中层融合策略整合眼底和OCT特征,并利用跨模态特征对齐模块减少模态差异,实现青光眼的精确分割与分级。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03521 2026-03-24 cs.MM cs.LG 83%

Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation

跨空间协同:一种用于对话中多模态情感识别的统一框架

Xiaosen Lyu, Jiayu Xiong, Yuren Chen, Wanlong Wang, Xiaoqing Dai, Jing Wang

机构 * Xiaosen Lyu 1,2(李绍森 1,2) Jiayu Xiong 1,2(熊佳宇 1,2) Yuren Chen 1,2(陈远人 1,2) Wanlong Wang 1,2(王万龙 1,2) Xiaoqing Dai 1,2(戴晓青 1,2) Jing Wang 1,2(王婧 1,2)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出Cross-Space Synergy框架,通过协同多项式融合和帕累托梯度调节器有效提升多模态情感识别的准确性和训练稳定性。

Comments Accepted to AAAI 2026

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(29), 24226-24234 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22862 2026-03-24 cs.LG cs.CV 83%

Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time Adaptation

通过渐进重对齐桥接模态以实现多模态测试时适应

Jiacheng Li, Songhe Feng

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出BriMPR框架,通过分治策略解决多模态测试时适应中的模态间分布偏移和语义对齐问题,通过提示调优和跨模态对比学习提升多模态特征对齐效果。

Comments Accepted by AAAI 2026 (Oral)

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence. 2026, 40(27): 22931-22939

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19623 2026-03-23 cs.CV 83%

Disentangle-then-Align: Non-Iterative Hybrid Multimodal Image Registration via Cross-Scale Feature Disentanglement

解耦后再对齐:通过跨尺度特征解耦实现非迭代混合多模态图像配准

Chunlei Zhang, Jiahao Xia, Yun Xiao, Bo Jiang, Jian Zhang

机构 * Faculty of Engineering and IT, University of Technology Sydney(新南威尔士大学工程与信息技术学院) School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出HRNet网络,通过解耦表示与混合参数预测,解决多模态图像配准中共享空间不稳定和单类型变换限制的问题,实现非迭代的粗到细配准。

Comments Accepted by CVPR 2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17885 2026-03-23 cs.CV cs.LG 83%

FastMMoE: Accelerating Multimodal Large Language Models through Dynamic Expert Activation and Routing-Aware Token Pruning

FastMMoE:通过动态专家激活和路由感知的标记剪枝加速多模态大语言模型

Guoyang Xia, Yifeng Ding, Fengfa Li, Lei Ren, Wei Chen, Fangxiang Feng, Xiaojie Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Li Auto(利亚自动化)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出FastMMoE,一种无需训练的加速框架,通过动态专家激活和路由感知标记剪枝,显著降低计算量并保持性能,优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19026 2026-03-20 cs.CV 83%

Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token

重新思考MLLM本身作为分割器:仅用一个分割标记

Anqi Zhang, Xiaokang Ji, Guangyu Gao, Jianbo Jiao, Chi Harold Liu, Yunchao Wei

机构 * Beijing Institute of Technology(北京理工大学) University of Birmingham(伯明翰大学) Beijing Jiaotong University(北京交通大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出SELF1E方法,通过保留原始图像分辨率并利用残差特征提升分割精度,无需外部解码器即可实现与专业解码器相当的分割性能。

Comments Paper is accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17705 2026-03-19 cs.CV 83%

Parameter-Efficient Modality-Balanced Symmetric Fusion for Multimodal Remote Sensing Semantic Segmentation

参数高效模态平衡对称融合用于多模态遥感语义分割

Haocheng Li, Juepeng Zheng, Shuangxi Miao, Ruibo Lu, Guosheng Cai, Haohuan Fu, Jianxi Huang

机构 * College of Land Science and Technology, China Agricultural University(中国农业大学土地科学与技术学院) Key Laboratory of Remote Sensing for Agri-Hazards, Ministry of Agriculture and Rural Affairs(农业农村部农业灾害遥感重点实验室) Faculty of Geosciences and Engineering, Southwest Jiaotong University(西南交通大学地质科学与工程学院) School of Artificial Intelligence, Sun Yat-Sen University(中山大学人工智能学院) Henan Polytechnic University(河南理工大学) Key Laboratory of Spatio-Temporal Information and Ecological Restoration of Mines, Ministry of Natural Resources of the People’s Republic of China(矿产资源时空信息与生态修复重点实验室) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) National Supercomputing Center in Shenzhen, Shenzhen, China(深圳国家超算中心) Ministry of Education Key Laboratory for Earth System Modeling and the Department of Earth System Science, Tsinghua University(地球系统模拟教育部重点实验室和清华大学地球系统科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MoBaNet,通过参数高效和模态平衡的对称融合框架,在减少可训练参数的同时提升多模态遥感语义分割的鲁棒性和平衡性。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16421 2026-03-19 cs.CV 83%

HGP-Mamba: Integrating Histology and Generated Protein Features for Mamba-based Multimodal Survival Risk Prediction

HGP-Mamba:整合组织学与生成的蛋白质特征用于基于Mamba的多模态生存风险预测

Jing Dai, Chen Wu, Ming Wu, Qibin Zhang, Zexi Wu, Jingdong Zhang, Hongming Xu

机构 * Cancer Hospital of Dalian University of Technology, Shenyang, China(大连理工大学沈阳医院) School of Biomedical Engineering, Faculty of Medicine, Dalian University of Technology, Dalian, China(大连理工大学生物医学工程学院) Key Laboratory of Integrated Circuit and Biomedical Electronic System, Dalian University of Technology, Dalian, China(大连理工大学集成电路与生物医学电子系统重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出HGP-Mamba框架,通过整合组织学与生成蛋白质特征,提升多模态生存风险预测的效率与性能。

Comments Accepted at IEEE ICME 2026. This arXiv version includes additional supplementary experiments and extended discussions beyond the conference version

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15818 2026-03-18 cs.CV 83%

Conflict-Aware Multimodal Fusion for Ambivalence and Hesitancy Recognition

具有冲突意识的多模态融合用于矛盾与犹豫识别

Salah Eddine Bekhouche, Hichem Telli, Azeddine Benlamoudi, Salah Eddine Herrouz, Abdelmalik Taleb-Ahmed, Abdenour Hadid

机构 * University of the Basque Country UPV/EHU(巴斯克大学UPV/EHU) Laboratory of LESIA, University of Biskra(贝斯克拉大学LESIA实验室) Lab. de Génie Electrique (LAGE), University Kasdi Merbah Ouargla(奥尔加拉大学LAGE实验室) Institute of Electronics, Microelectronics and Nanotechnology (IEMN), Polytechnic University of Hauts-de-France, University of Lille(电子、微电子与纳米技术研究所(IEMN),法国 Hauts-de-France 工业大学,里尔大学) Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi, UAE(索邦人工智能中心,索邦大学阿布扎比校区,阿联酋)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出ConflictAwareAH框架,通过多模态融合识别矛盾与犹豫状态,提升F1指标,采用冲突特征作为双向线索,改进模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14132 2026-03-17 cs.CV cs.LG 83%

DualSwinFusionSeg: Multimodal Martian Landslide Segmentation via Dual Swin Transformer with Multi-Scale Fusion and UNet++

DualSwinFusionSeg: 多模态火星滑坡分割 via 双Swin Transformer 与多尺度融合及UNet++

Shahriar Kabir, Abdullah Muhammed Amimul Ehsan, Istiak Ahmmed Rifti, Md Kaykobad Reza

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出DualSwinFusionSeg,通过双Swin Transformer与多尺度融合及UNet++实现多模态火星滑坡分割,验证了在有限标注数据下提升分割精度的效果。

Comments 10 pages, 2 Figures, 12 Tables. Code is available at: https://github.com/amimulamim/Mars-LS-Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13719 2026-03-17 cs.CV 83%

Sparse-Dense Mixture of Experts Adapter for Multi-Modal Tracking

稀疏-密集专家混合适配器用于多模态跟踪

Yabin Zhu, Jianqi Li, Chenglong Li, Jiaxiang Wang, Chengjie Gu, Jin Tang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出稀疏-密集专家混合适配器框架,通过稀疏MoE和密集共享MoE有效建模多模态特征,结合Gram基于语义对齐超图融合模块,提升多模态跟踪性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13272 2026-03-17 cs.LG cs.AI 83%

CAMEL-CLIP: Channel-aware Multimodal Electroencephalography-text Alignment for Generalizable Brain Foundation Models

CAMEL-CLIP:面向通用脑基础模型的通道感知多模态EEG-文本对齐

Hanseul Choi, Jinyeong Park, Seongwon Jin, Sungho Park, Jibum Kim

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Incheon National University(庆尚国立大学) Department of Artificial Intelligence(人工智能系) Inha University(Inha大学) Center for Brain-Machine Interface(脑机接口中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 CAMEL-CLIP通过通道属性位置编码、动态通道投影和双级对比学习,提升EEG-文本多模态基础模型在通道异质性下的鲁棒性,实验表明其在线性探测中表现最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05391 2026-03-13 cs.CV 83%

LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models

LoC-Path:学习压缩用于病理多模态大语言模型

Qingqiao Hu, Weimin Lyu, Meilong Xu, Kehan Qi, Xiaoling Hu, Saumya Gupta, Jiawei Zhou, Chao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 LoC-Path通过压缩滑片图像特征,降低多模态模型的训练和推理成本,使在有限资源下实现高效病理大语言模型成为可能。

Comments Code will be released soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11306 2026-03-13 cs.CV 83%

Hierarchical Granularity Alignment and State Space Modeling for Robust Multimodal AU Detection in the Wild

层次粒度对齐与状态空间建模用于野外多模态面部动作单元检测

Jun Yu, Yunxiang Zhang, Naixiang Zheng, Lingsi Zhu, Guoyuan Wang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

AI总结 本文提出了一种基于层次粒度对齐和状态空间模型的多模态框架,通过强大的基础模型提取高保真视觉和音频表示,并引入视觉-马尔可夫模型和不对称交叉注意机制,实现对野外环境中面部动作单元的高效检测。

Comments 8 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10877 2026-03-12 cs.CL 83%

From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers

从图像到词语:高效的跨模态知识蒸馏用于从黑盒教师模型向语言模型转移

Ayan Sengupta, Shantanu Dixit, Md Shad Akhtar, Tanmoy Chakraborty

机构 * Indian Institute of Technology Delhi, India(印度德里印度理工学院) Indraprastha Institute of Information Technology Delhi, India(印度德里印度信息科技学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL

AI总结 本文提出ARMADA框架,通过高效的跨模态知识蒸馏方法,从黑盒视觉-语言模型向语言模型转移知识,实现性能提升且无需昂贵预训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10370 2026-03-12 cs.CV 83%

GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning

GeoSense: 内化几何必要性感知以实现多模态推理

Ruiheng Liu, Haihong Hao, Mingfei Han, Xin Gu, Kecheng Zhang, Changlin Li, Xiaojun Chang

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 GeoSense通过内化几何必要性感知,提升多模态推理的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09173 2026-03-11 cs.CV 83%

Point Cloud as a Foreign Language for Multi-modal Large Language Model

点云作为多模态大语言模型的外语

Sneha Paul, Zachary Patterson, Nizar Bouguila

机构 * Concordia University(康科迪亚大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 SAGE是首个端到端3D多模态大语言模型,通过轻量级3D分词器直接处理点云,提升3D任务的推理能力与鲁棒性。

Comments Accepted in The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09111 2026-03-11 cs.CV 83%

Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities

逐步表示学习用于不完整模态的多模态情感分析

Jindi Bao, Jianjun Qian, Mengkai Yan, Jian Yang

机构 * Nanjing University of Science and Technology(南京理工大学) Hohai University(河海大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 PRLF通过逐步表示学习框架,在不完整模态条件下提升多模态情感分析的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08800 2026-03-11 cs.CV 83%

Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM

Granulon: 通过自适应多粒度语义唤醒像素级视觉编码器以实现MLLM

Junyuan Mao, Qiankun Li, Linghao Meng, Zhicheng He, Xinliang Zhou, Kun Wang, Yang Liu, Yueming Jin

机构 * National University of Singapore(国立新加坡大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 Granulon通过自适应多粒度语义增强,提升多模态大语言模型的像素级视觉理解和多粒度推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04320 2026-03-05 cs.IR cs.MM 83%

CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential Recommendation

CAMMSR: 基于类别的注意力混合专家多模态序列推荐

Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Hewei Wang, Yijie Li, Jianheng Tang, Yunhuai Liu, Edith C. H. Ngai

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 CAMMSR通过引入基于类别的注意力混合专家模块,实现多模态序列推荐的自适应、协同和用户中心化

Comments Accepted by ICDE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03827 2026-03-05 cs.MM 83%

Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition

通过层次语义表示进行进化多模态推理以实现意图识别

Qianrui Zhou, Hua Xu, Yunjin Gu, Yifan Wang, Songze Li, Hanlei Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.MM

AI总结 HIER通过层次语义表示与进化推理提升多模态意图识别性能,实现更准确的意图推断。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03544 2026-03-05 cs.CV 83%

PinCLIP: Large-scale Foundational Multimodal Representation at Pinterest

PinCLIP:Pinterest的规模化多模态基础表示

Josh Beal, Eric Kim, Jinfeng Rao, Rex Wu, Dmitry Kislyuk, Charles Rosenberg

机构 * Pinterest Inc.(Pinterest公司)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 PinCLIP通过多模态表示学习提升Pinterest的检索和排序模型,解决冷启动问题并提高用户参与度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03276 2026-03-04 cs.CV 83%

Beyond Language Modeling: An Exploration of Multimodal Pretraining

超越语言模型:多模态预训练的探索

Shengbang Tong, David Fan, John Nguyen, Ellis Brown, Gaoyue Zhou, Shengyi Qian, Boyang Zheng, Théophane Vallaeys, Junlin Han, Rob Fergus, Naila Murray, Marjan Ghazvininejad, Mike Lewis, Nicolas Ballas, Amir Bar, Michael Rabbat, Jakob Verbeek, Luke Zettlemoyer, Koustuv Sinha, Yann LeCun, Saining Xie

机构 * FAIR, Meta(FAIR、Meta) New York University(纽约大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 本文通过多模态预训练探索,揭示了视觉与语言数据的互补性及统一预训练对世界建模的促进作用,并提出MoE架构解决多模态扩展的不对称性问题。

Comments Project website at https://beyond-llms.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02532 2026-03-04 cs.CV 83%

EIMC: Efficient Instance-aware Multi-modal Collaborative Perception

EIMC: 高效实例感知多模态协作感知

Kang Yang, Peng Wang, Lantao Li, Tianci Bu, Chen Sun, Deying Li, Yongcai Wang

机构 * School of Information, Renmin University of China(中国人民大学信息学院) Sony Research and Development Center China(索尼(中国)研发有限公司) National University of Defense Technology(国防科技大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 EIMC通过实例感知的多模态协作感知方法,提升自动驾驶安全性,减少带宽使用,实现高效且准确的3D感知。

Comments 9 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02505 2026-03-04 cs.CV 83%

SGMA: Semantic-Guided Modality-Aware Segmentation for Remote Sensing with Incomplete Multimodal Data

SGMA: 基于不完整多模态数据的语义引导模态感知遥感语义分割

Lekang Wen, Liang Liao, Jing Xiao, Mi Wang

机构 * State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(武汉大学测绘遥感信息工程国家重点实验室) Hangzhou Institute of Technology, Xidian University(西安电子科技大学杭州研究院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 SGMA通过语义引导和模态感知方法解决不完整多模态数据下的语义分割问题,提升模态间平衡性和一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02474 2026-03-04 cs.IR cs.AI 83%

Q-BERT4Rec: Quantized Semantic-ID Representation Learning for Multimodal Recommendation

Q-BERT4Rec: 量化语义-ID表示学习用于多模态推荐

Haofeng Huang, Ling Gai

机构 * University of Shanghai for Science and Technology(上海科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 Q-BERT4Rec通过统一语义表示与量化建模,提升多模态推荐的性能和可解释性。

Comments Submitted to KDD2026

详情

展开后加载摘要…

URL PDF HTML 收藏