arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2107.11853 2021-07-27 cs.CV 74%

Will Multi-modal Data Improves Few-shot Learning?

Zilun Zhang, Shihao Ma, Yichun Zhang

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Project Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.04288 2021-07-12 eess.IV cs.CV 74%

Retinal OCT Denoising with Pseudo-Multimodal Fusion Network

Dewei Hu, Joseph D. Malone, Yigit Atay, Yuankai K. Tao, Ipek Oguz

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments Accepted by International Workshop on Ophthalmic Medical Image Analysis (OMIA) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02358 2021-07-06 cs.CV cs.LG 74%

VisualWordGrid: Information Extraction From Scanned Documents Using A Multimodal Approach

Mohamed Kerroumi, Othmane Sayem, Aymen Shabou

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.13536 2021-05-31 eess.SP cs.CV cs.LG 74%

ECG Heart-beat Classification Using Multimodal Image Fusion

Zeeshan Ahmad, Anika Tabassum, Naimul Khan, Ling Guan

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.07803 2021-04-19 cs.CV 74%

Semisupervised Manifold Alignment of Multimodal Remote Sensing Images

Devis Tuia, Michele Volpi, Maxime Trolliet, Gustau Camps-Valls

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Journal ref IEEE Transactions on Geoscience and Remote Sensing, 52(12): 7708 - 7720, 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.07517 2020-11-17 cs.CV 74%

Data-efficient Alignment of Multimodal Sequences by Aligning Gradient Updates and Internal Feature Distributions

Jianan Wang, Boyang Li, Xiangyu Fan, Jing Lin, Yanwei Fu

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments This paper is accepted to WACV2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.01935 2020-11-03 cs.RO cs.AI cs.LG cs.SY eess.SY 74%

Probabilistic End-to-End Vehicle Navigation in Complex Dynamic Environments with Multimodal Sensor Fusion

Peide Cai, Sukai Wang, Yuxiang Sun, Ming Liu

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

Comments 8 pages, 6 figures, 3 tables. IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.10991 2020-10-13 eess.AS cs.HC cs.LG stat.ML 74%

Attention Driven Fusion for Multi-Modal Emotion Recognition

Darshana Priyasad, Tharindu Fernando, Simon Denman, Clinton Fookes, Sridha Sridharan

专题命中 多模态训练与对齐 :multi-modal(title);分类 eess.AS

Comments An updated version of the ICASSP 2020 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.16607 2020-07-01 cs.NE cs.CV cs.LG 74%

A Framework for Learning Invariant Physical Relations in Multimodal Sensory Processing

Du Xiaorui, Yavuzhan Erdem, Immanuel Schweizer, Cristian Axenie

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.10738 2020-06-23 cs.CV 74%

Non-linear and Selective Fusion of Cross-Modal Images

Aiqing Fang, Xinbo Zhao, Jiaqi Yang, Yanning Zhang

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08577 2020-06-23 cs.CV cs.IT cs.LG math.IT 74%

A Cross-Modal Image Fusion Method Guided by Human Visual Characteristics

Aiqing Fang, Xinbo Zhao, Jiaqi Yang, Yanning Zhang

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.06156 2020-04-10 cs.CV cs.RO 74%

Gimme Signals: Discriminative signal encoding for multimodal activity recognition

Raphael Memmesheimer, Nick Theisen, Dietrich Paulus

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments 8 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.01131 2020-03-31 cs.CL cs.SI 74%

See and Read: Detecting Depression Symptoms in Higher Education Students Using Multimodal Social Media Data

Paulo Mann, Aline Paes, Elton H. Matsushima

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL

Comments This article was accepted (15 November 2019) and will appear in the proceedings of ICWSM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.07753 2020-02-21 cs.CV 74%

Multimodal feature fusion for CNN-based gait recognition: an empirical comparison

Francisco Manuel Castro, Manuel Jesús Marín-Jiménez, Nicolás Guil, Nicolás Pérez de la Blanca

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:1603.01006

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.00216 2020-02-04 cs.RO cs.CV 74%

Leveraging Uncertainties for Deep Multi-modal Object Detection in Autonomous Driving

Di Feng, Yifan Cao, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.10718 2019-12-24 cs.CV 74%

Cross-Modal Image Fusion Theory Guided by Subjective Visual Attention

Aiqing Fang, Xinbo Zhao, Yanning Zhang

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08487 2019-12-20 cs.CV cs.RO 74%

FuseSeg: LiDAR Point Cloud Segmentation Fusing Multi-Modal Data

Georg Krispel, Michael Opitz, Georg Waltner, Horst Possegger, Horst Bischof

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Accepted for publication in WACV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.11714 2019-09-02 cs.CV 74%

Multi-Modal Fusion for End-to-End RGB-T Tracking

Lichao Zhang, Martin Danelljan, Abel Gonzalez-Garcia, Joost van de Weijer, Fahad Shahbaz Khan

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Accepted at ICCVW (VOT) 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.07754 2019-05-21 cs.CV cs.LG 74%

Multimodal 3D Object Detection from Simulated Pretraining

Åsmund Brekke, Fredrik Vatsendvik, Frank Lindseth

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments 12 pages, part of proceedings for the NAIS 2019 symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.10796 2019-03-07 cs.CV cs.CY 74%

Dynamic Deep Multi-modal Fusion for Image Privacy Prediction

Ashwini Tonge, Cornelia Caragea

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Accepted by The Web Conference (WWW) 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.11155 2017-12-01 cs.CV 74%

Predicting Depression Severity by Multi-Modal Feature Engineering and Fusion

Aven Samareh, Yan Jin, Zhangyang Wang, Xiangyu Chang, Shuai Huang

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-18)

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.10422 2017-11-02 cs.RO cs.AI cs.LG 74%

Learning End-to-end Multimodal Sensor Policies for Autonomous Navigation

Guan-Horng Liu, Avinash Siravuru, Sai Prabhakar, Manuela Veloso, George Kantor

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

Comments to be published in Conference on Robot Learning (CoRL), 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.01266 2017-04-06 cs.CV cs.CY cs.HC 74%

Supporting Navigation of Outdoor Shopping Complexes for Visually-impaired Users through Multi-modal Data Fusion

Archana Paladugu, Parag S. Chandakkar, Peng Zhang, Baoxin Li

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments ICME 2013

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.04117 2016-03-15 cs.CV 74%

Multi-modal Tracking for Object based SLAM

Prateek Singhal, Ruffin White, Henrik Christensen

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Submitted to IROS 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06411 2026-08-10 cs.AI cs.CV 新提交 73%

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

学习在多模态大语言模型(MLLMs)中预测中间层注意力以进行视觉token剪枝

Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang, Hao Geng, Minjun Yu

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 该研究针对MLLMs视觉token剪枝的固定层非最优及计算成本高的问题,提出MAP方法,实现仅保留5.56%视觉token时维持97.5%性能,获3.09倍端到端加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03247 2026-08-05 cs.CV cs.CL 新提交 73%

CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment

CIGTSurv:结合局部原型关联与全局特征对齐的临床信息引导三模态生存预测

Jing Dai, Qibin Zhang, Weiwei Zhou, Mingde Xu, Jingsong Liu, Jingdong Zhang, Hongming Xu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本研究针对临床信息未充分利用及多模态异质性问题,提出CIGTSurv框架,结合局部原型关联与全局特征对齐机制,在五个TCGA癌症队列上取得生存预测SOTA性能。

Comments Accepted at MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22712 2026-07-29 cs.CV cs.AI physics.optics 版本更新 73%

scMIR: a vision-language foundation model for single-cell light microscopy image representation

scMIR:用于单细胞光学显微镜图像表示的视觉语言基础模型

Yifan Shang, Jiahui Tan, Xiangxiang Zeng, Renjie Zhou

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 研究针对单细胞光学显微镜图像分析难题,提出scMIR模型,通过自监督图像重建与文本引导跨模态对齐,在多图像文本对上预训练,在多种复杂任务中表现出色,优于现有方法,具备强泛化能力,能推动高通量表型分析工作流程标准化和自动化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09481 2026-07-13 cs.CV cs.AI 新提交 73%

Decoupling Language Guidance from Backbones for Text-Guided Medical Segmentation

用于文本引导医学分割的骨干网络与语言引导解耦

Yungeng Liu, Xuanzi Fang, Haijin Zeng, Qi Dai, Yongyong Chen

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) NingBo No.2 Hospital(宁波第二医院)

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 研究针对文本引导医学分割中模型组件紧密耦合问题,提出可转移骨干层次适配器框架BTHA,通过稳定特征级接口、分层监督策略和自适应门控语义引导适配器,有效提升分割效果且计算开销适度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06620 2026-07-09 cs.CV cs.AI 新提交 73%

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts

SpaR3D-MoE:来自稀疏视图的自适应3D空间推理与几何归纳专家混合模型

Haida Feng, Hao Wei, Haolin Wang, Shiwei Li, Chade Li, Yihong Wu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) YUKUN Intelligent World(宇琨智能世界)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 研究针对多模态大语言模型在2D与3D表征差距问题,提出SpaR3D-MoE框架,通过自适应时空采样和几何归纳专家混合模型,从稀疏RGB输入实现自适应空间推理,在多个实验中取得最优性能。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02484 2026-07-03 cs.CV cs.AI 新提交 73%

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning

对抗文本噪声与冗余:熵感知的密集视觉令牌剪枝

Xuehui Wang, Xuankun Yang, Wei Shen

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出熵感知密集剪枝(EADP)框架,通过熵过滤文本噪声并利用子模最大化选择令牌,在严格预算下保留细粒度视觉线索,提升VLM精度-效率权衡。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏