arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2504.06543 2025-04-10 cs.IR 78%

DiffusionCom: Structure-Aware Multimodal Diffusion Model for Multimodal Knowledge Graph Completion

Wei Huang, Meiyu Liang, Peining Li, Xu Hou, Yawen Li, Junping Du, Zhe Xue, Zeli Guan

专题命中 多模态生成 :multimodal(title,abstract)

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04209 2025-04-01 cs.RO 78%

CALMM-Drive: Confidence-Aware Autonomous Driving with Large Multimodal Model

Ruoyu Yao, Yubin Wang, Haichao Liu, Rui Yang, Zengqi Peng, Lei Zhu, Jun Ma

专题命中 多模态生成 :multimodal(title,abstract)

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20913 2025-03-28 cs.CE cs.LG 78%

TransDiffSBDD: Causality-Aware Multi-Modal Structure-Based Drug Design

Xiuyuan Hu, Guoqing Liu, Can Chen, Yang Zhao, Hao Zhang, Xue Liu

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07975 2025-03-25 cs.CV cs.AI cs.CL 78%

JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Haowei Zhang, Xingkai yu, Liang Zhao, Yisong Wang, Jiaying Liu, Chong Ruan

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08937 2025-03-13 eess.SP cs.LG 78%

Beam Selection in ISAC using Contextual Bandit with Multi-modal Transformer and Transfer Learning

Mohammad Farzanullah, Han Zhang, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 6 pages, 4 figures, 2 tables, IEEE International Conference on Communications 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06119 2025-03-11 cs.LG 78%

Unlocking Pretrained LLMs for Motion-Related Multimodal Generation: A Fine-Tuning Approach to Unify Diffusion and Next-Token Prediction

Shinichi Tanaka, Zhao Wang, Yoichi Kato, Jun Ohya

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03752 2025-03-07 cs.CY 78%

Multimodal Generative AI and Foundation Models for Behavioural Health in Online Gambling

Konrad Samsel, Mohammad Noaeen, Neil Seeman, Karim Keshavjee, Li-Jia Li, Zahra Shakeri

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11734 2025-03-04 q-bio.QM cs.LG q-bio.GN 78%

Multi-Modal and Multi-Attribute Generation of Single Cells with CFGen

Alessandro Palma, Till Richter, Hanyi Zhang, Manuel Lubetzki, Alexander Tong, Andrea Dittadi, Fabian Theis

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 41 pages, 22 figures

Journal ref The Thirteenth International Conference on Learning Representations (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11399 2025-02-19 cs.HC 78%

FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization

Yuki Tatsukawa, I-Chao Shen, Mustafa Doga Dogan, Anran Qi, Yuki Koyama, Ariel Shamir, Takeo Igarashi

专题命中 多模态生成 :multimodal(title,abstract)

Comments 14 pages

Journal ref CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23402 2025-02-11 cs.SE 78%

VisualCoder: Guiding Large Language Models in Code Execution with Fine-grained Multimodal Chain-of-Thought Reasoning

Cuong Chi Le, Hoang-Chau Truong-Vinh, Huy Nhat Phan, Dung Duy Le, Tien N. Nguyen, Nghi D. Q. Bui

专题命中 多模态生成 :multimodal(title,abstract)

Comments NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12849 2025-01-31 cs.RO 78%

Whole-Body Trajectory Optimization for Robot Multimodal Locomotion

Giuseppe L'Erario, Gabriele Nava, Giulio Romualdi, Fabio Bergonti, Valentino Razza, Stefano Dafarra, Daniele Pucci

专题命中 多模态生成 :multimodal(title,abstract)

Comments Paper accepted in Humanoids 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08825 2025-01-16 eess.SP 78%

A Multi-modal Intelligent Channel Model for 6G Multi-UAV-to-Multi-Vehicle Communications

Lu Bai, Mengyuan Lu, Ziwei Huang, Xiang Cheng

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07333 2025-01-14 eess.SP 78%

Synesthesia of Machines Based Multi-Modal Intelligent V2V Channel Model

Zengrui Han, Lu Bai, Ziwei Huang, Xiang Cheng

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12933 2025-01-14 cs.SE 78%

ACTesting: Automated Cross-modal Testing Method of Text-to-Image Software

Siqi Gu, Chunrong Fang, Quanjun Zhang, Zhenyu Chen

专题命中 多模态生成 :cross-modal(title,abstract)

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17420 2024-12-03 cs.CE eess.IV 78%

Cross-modal Medical Image Generation Based on Pyramid Convolutional Attention Network

Fuyou Mao, Lixin Lin, Ming Jiang, Dong Dai, Chao Yang, Hao Zhang, Yan Tang

专题命中 多模态生成 :cross-modal(title);multimodal(abstract)

Comments 18 pages, 6 figures, Machine Vision and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15220 2024-11-26 cs.LG cs.NA math.NA stat.CO stat.ML 78%

Sampling with Adaptive Variance for Multimodal Distributions

Björn Engquist, Kui Ren, Yunan Yang

专题命中 多模态生成 :multimodal(title,abstract)

Comments 26 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11394 2024-11-19 cs.RO 78%

InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models

Yu Yan, Rongtao Xu, Jiazhao Zhang, Peiyang Li, Xiaodan Liang, Jianqin Yin

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09117 2024-11-15 cs.LG cs.DS math.PR stat.ML 78%

Efficiently learning and sampling multimodal distributions with data-based initialization

Frederic Koehler, Holden Lee, Thuy-Duong Vuong

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03711 2024-11-07 eess.SP 78%

Multi-Modal Intelligent Channel Modeling: A New Modeling Paradigm via Synesthesia of Machines

Lu Bai, Ziwei Huang, Mingran Sun, Xiang Cheng, Lizhen Cui

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.04003 2024-10-30 eess.SP 78%

Modeling Time-dependent CO$_2$ Intensities in Multi-modal Energy Systems with Storage

Christopher Ripp, Florian Steinke

专题命中 多模态生成 :multi-modal(title,abstract)

Comments This work has been submitted to the Elsevier Applied Energy for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13782 2024-10-18 cs.LG q-bio.QM 78%

DPLM-2: A Multimodal Diffusion Protein Language Model

Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, Quanquan Gu

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10347 2024-10-10 cs.IT math.IT 78%

Editable-DeepSC: Cross-Modal Editable Semantic Communication Systems

Wenbo Yu, Bin Chen, Qinshan Zhang, Shu-Tao Xia

专题命中 多模态生成 :cross-modal(title,abstract)

Comments published at VTC2024-Spring

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00712 2024-10-04 q-bio.NC cs.LG 78%

NECOMIMI: Neural-Cognitive Multimodal EEG-informed Image Generation with Diffusion Models

Chi-Sheng Chen

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08790 2024-10-04 eess.IV 78%

A Multimodal Approach for Fluid Overload Prediction: Integrating Lung Ultrasound and Clinical Data

Tianqi Yang, Nantheera Anantrasirichai, Oktay Karakuş, Marco Allinovi, Alin Achim

专题命中 多模态生成 :multimodal(title,abstract)

Comments In the experiment, for the classification tasks, the network was informed with ground truth during training, significantly improving the performance. This makes the results invalid. Therefore, corrections and more validations are needed to evaluate the performance of the method

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17506 2024-09-27 cs.NI 78%

Optimizing Resource Allocation for Multi-modal Semantic Communication in Mobile AIGC Networks: A Diffusion-based Game Approach

Jian Liu, Ming Xiao, Jinbo Wen, Jiawen Kang, Ruichen Zhang, Tao Zhang, Dusit Niyato, Weiting Zhang, Ying Liu

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08269 2024-09-13 cs.RO 78%

Touch2Touch: Cross-Modal Tactile Generation for Object Manipulation

Samanta Rodriguez, Yiming Dou, Miquel Oller, Andrew Owens, Nima Fazeli

专题命中 多模态生成 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08078 2024-08-20 cs.CV cs.AI cs.CL 78%

Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation

Wenting Chen, Linlin Shen, Jingyang Lin, Jiebo Luo, Xiang Li, Yixuan Yuan

专题命中 多模态生成 :image-text(title);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACL 2024

Journal ref https://aclanthology.org/2024.acl-long.514/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04423 2024-08-09 cs.RO 78%

UNMuTe: Unifying Navigation and Multimodal Dialogue-like Text Generation

Niyati Rawal, Roberto Bigazzi, Lorenzo Baraldi, Rita Cucchiara

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15132 2024-07-23 q-bio.NC cs.LG 78%

Deep multimodal saliency parcellation of cerebellar pathways: linking microstructure and individual function through explainable multitask learning

Ari Tchetchenian, Leo Zekelman, Yuqian Chen, Jarrett Rushmore, Fan Zhang, Edward H. Yeterian, Nikos Makris, Yogesh Rathi, Erik Meijering, Yang Song, Lauren J. O'Donnell

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06947 2024-07-18 eess.SP cs.LG cs.NI 78%

A Multi-Modal Simulation Framework to Enable Digital Twin-based V2X Communications in Dynamic Environments

Lorenzo Cazzella, Francesco Linsalata, Maurizio Magarini, Matteo Matteucci, Umberto Spagnolini

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏