arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2502.07391 2025-02-12 cs.CL 79%

Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation Generation

Palaash Goel, Dushyant Singh Chauhan, Md Shad Akhtar

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06583 2025-02-11 cs.CV 79%

Adaptive Perception for Unified Visual Multi-modal Object Tracking

Xiantao Hu, Bineng Zhong, Qihua Liang, Zhiyi Mo, Liangtao Shi, Ying Tai, Jian Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06100 2025-02-11 cs.CV eess.SP 79%

Col-OLHTR: A Novel Framework for Multimodal Online Handwritten Text Recognition

Chenyu Liu, Jinshui Hu, Baocai Yin, Jia Pan, Bing Yin, Jun Du, Qingfeng Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06062 2025-02-11 eess.IV cs.AI 79%

Multi-modal Data Fusion and Deep Ensemble Learning for Accurate Crop Yield Prediction

Akshay Dagadu Yewle, Laman Mirzayeva, Oktay Karakuş

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 28 pages, 7 figures and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04220 2025-02-10 cs.RO cs.AI cs.LG 79%

Multi-modal perception for soft robotic interactions using generative models

Enrico Donato, Egidio Falotico, Thomas George Thuruthel

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.AI

Comments Accepted for presentation at IEEE RoboSoft 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18407 2025-02-10 cs.LG cs.AI q-bio.BM 79%

A Group Symmetric Stochastic Differential Equation Model for Molecule Multi-modal Pretraining

Shengchao Liu, Weitao Du, Zhiming Ma, Hongyu Guo, Jian Tang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14455 2025-02-07 cs.CV 79%

Triple Path Enhanced Neural Architecture Search for Multimodal Fake News Detection

Bo Xu, Qiujie Xie, Jiahui Zhou, Linlin Zong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments IEEE International Conference on Acoustics, Speech, and Signal Processing(ICASSP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02938 2025-02-06 cs.CL 79%

LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier

T. Chay-intr, Y. Chen, K. Viriyayudhakorn, T. Theeramunkong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08167 2025-02-05 cs.LG cs.CL q-bio.QM 79%

MolBind: Multimodal Alignment of Language, Molecules, and Proteins

Teng Xiao, Chao Cui, Huaisheng Zhu, Vasant G. Honavar

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01467 2025-02-04 cs.CV 79%

Deep Unfolding Multi-modal Image Fusion Network via Attribution Analysis

Haowen Bai, Zixiang Zhao, Jiangshe Zhang, Baisong Jiang, Lilun Deng, Yukun Cui, Shuang Xu, Chunxia Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted in IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00078 2025-02-04 eess.IV cs.CV 79%

Deep Ensembling with Multimodal Image Fusion for Efficient Classification of Lung Cancer

Surochita Pal, Sushmita Mitra

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05040 2025-02-04 cs.CV 79%

Unsupervised Multimodal 3D Medical Image Registration with Multilevel Correlation Balanced Optimization

Jiazheng Wang, Xiang Chen, Yuxi Zhang, Min Liu, Yaonan Wang, Hang Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Method description for MICCAI Learn2Reg 2024 challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15099 2025-01-28 cs.CV cs.LG 79%

Bringing RGB and IR Together: Hierarchical Multi-Modal Enhancement for Robust Transmission Line Detection

Shengdong Zhang, Xiaoqin Zhang, Wenqi Ren, Linlin Shen, Shaohua Wan, Jun Zhang, Yujing M Jiang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10791 2025-01-28 cs.CV 79%

CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes

Tim Broedermann, Christos Sakaridis, Yuqian Fu, Luc Van Gool

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments IEEE Robotics and Automation Letters, The source code is publicly available at: https://github.com/timbroed/CAFuser

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09323 2025-01-28 cs.CV 79%

E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection

Jiaqing Zhang, Mingxiang Cao, Weiying Xie, Jie Lei, Daixun Li, Wenbo Huang, Yunsong Li, Xue Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11968 2025-01-22 cs.AI cs.LG 79%

Bridging Visualization and Optimization: Multimodal Large Language Models on Graph-Structured Combinatorial Optimization

Jie Zhao, Kang Hao Cheong, Witold Pedrycz

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10958 2025-01-22 cs.CV 79%

Rethinking Early-Fusion Strategies for Improved Multimodal Image Segmentation

Zhengwen Shen, Yulian Li, Han Zhang, Yuchen Weng, Jun Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00320 2025-01-22 cs.AI 79%

Advancing Multimodal Data Fusion in Pain Recognition: A Strategy Leveraging Statistical Correlation and Human-Centered Perspectives

Xingrui Gu, Zhixuan Wang, Irisa Jin, Zekun Wu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by AHRI 2024

Journal ref 2024 12th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.07289 2025-01-22 cs.CV eess.IV 79%

VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization

Kaichen Zhou, Changhao Chen, Bing Wang, Muhamad Risqi U. Saputra, Niki Trigoni, Andrew Markham

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Journal ref The Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19070 2025-01-20 cs.CV 79%

Myriad: Large Multimodal Model by Applying Vision Experts for Industrial Anomaly Detection

Yuanze Li, Haolin Wang, Shihao Yuan, Ming Liu, Debin Zhao, Yiwen Guo, Chen Xu, Guangming Shi, Wangmeng Zuo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12636 2025-01-20 cs.LG cs.AI eess.SP 79%

Streamlining Multimodal Data Fusion in Wireless Communication and Sensor Networks

Mohammud J. Bocus, Xiaoyang Wang, Robert. J. Piechocki

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 12 figures, 3 tables, under review in IEEE Transactions on Cognitive Communications and Networking

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08490 2025-01-16 cs.CV cs.LG 79%

FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing

Isaac Corley, Simone Fobi Nsutezo, Anthony Ortiz, Caleb Robinson, Rahul Dodhia, Juan M. Lavista Ferres, Peyman Najafirad

专题命中 多模态训练与对齐 :multimodal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07166 2025-01-14 cs.AI 79%

Natural Language-Assisted Multi-modal Medication Recommendation

Jie Tan, Yu Rong, Kangfei Zhao, Tian Bian, Tingyang Xu, Junzhou Huang, Hong Cheng, Helen Meng

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments 10 pages

Journal ref Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, Boise, ID, USA, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05993 2025-01-13 cs.CV 79%

Aria: An Open Multimodal Native Mixture-of-Experts Model

Dongxu Li, Yudong Liu, Haoning Wu, Yue Wang, Zhiqi Shen, Bowen Qu, Xinyao Niu, Fan Zhou, Chengen Huang, Yanpeng Li, Chongyan Zhu, Xiaoyi Ren, Chao Li, Yifan Ye, Peng Liu, Lihuan Zhang, Hanshu Yan, Guoyin Wang, Bei Chen, Junnan Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04373 2025-01-09 cs.CV 79%

FGU3R: Fine-Grained Fusion via Unified 3D Representation for Multimodal 3D Object Detection

Guoxin Zhang, Ziying Song, Lin Liu, Zhonghong Ou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15209 2025-01-09 cs.CV 79%

MSCoTDet: Language-driven Multi-modal Fusion for Improved Multispectral Pedestrian Detection

Taeheon Kim, Sangyun Chung, Damin Yeom, Youngjoon Yu, Hak Gu Kim, Yong Man Ro

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03332 2025-01-08 cs.CV 79%

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets

Tanay Agrawal, Mohammed Guermal, Michal Balazia, Francois Bremond

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Preprint. Final paper accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, February, 2025. 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05300 2025-01-08 cs.LG cs.AI 79%

Unity by Diversity: Improved Representation Learning in Multimodal VAEs

Thomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard, Norbert Fortin, Julia E. Vogt, Babak Shahbaba, Stephan Mandt

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at Neurips 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20280 2025-01-03 cs.LG cs.AI 79%

Sparsely Multimodal Data Fusion

Josiah Bjorgaard

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20487 2024-12-31 cs.LG cs.CV cs.IT math.IT 79%

Multimodal Variational Autoencoder: a Barycentric View

Peijie Qiu, Wenhui Zhu, Sayantan Kumar, Xiwen Chen, Xiaotong Sun, Jin Yang, Abolfazl Razi, Yalin Wang, Aristeidis Sotiras

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏