arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2506.22274 2026-02-11 cs.CV cs.CL 76%

Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication

常见物体脱离上下文(COOCo):研究多模态上下文和语义场景违规在指代通信中的作用

Filippo Merlo, Ece Takmaz, Wenkai Chen, Albert Gatt

机构 * Utrecht University(乌特雷赫大学) University of Trento(特伦托大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL

AI总结 COOCo研究了VLMs在指代生成中如何利用场景上下文,发现模型根据语义相关性和噪声水平动态平衡局部与上下文信息。

Comments Accepted to TACL (pre-MIT Press publication version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15392 2026-01-23 cs.AI cs.CV cs.LG 76%

GeMM-GAN: A Multimodal Generative Model Conditioned on Histopathology Images and Clinical Descriptions for Gene Expression Profile Generation

GeMM-GAN: 一种基于组织病理图像和临床描述的多模态生成模型,用于基因表达谱生成

Francesca Pia Panaccione, Carlo Sgaravatti, Pietro Pinoli

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

AI总结 GeMM-GAN通过结合图像和文本信息生成逼真的基因表达谱,提升疾病预测准确性。

Comments 12 pages, 2 figures. Published at Image Analysis and Processing - ICIAP 2025 Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13729 2025-12-17 cs.LG cs.AI cs.CV 76%

Composite Classifier-Free Guidance for Multi-Modal Conditioning in Wind Dynamics Super-Resolution

多模态条件下的风动力超分辨率复合分类器引导方法

Jacob Schnell, Aditya Makkar, Gunadi Gani, Aniket Srinivasan Ashok, Darren Lo, Mike Optis, Alexander Wong, Yuhao Chen

机构 * University of Waterloo(滑铁卢大学) Veer Renewables

专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.AI

AI总结 本文提出复合分类器引导方法,用于多模态条件下的风动力超分辨率重建,实现高保真度与低成本的风数据生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13689 2025-11-19 cs.CL cs.CV 76%

Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation

Sofia Jamil, Kotla Sai Charan, Sriparna Saha, Koustava Goswami, Joseph K J

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02046 2025-11-05 cs.CV cs.AI 76%

Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis

Soham Joshi, Shwet Kamal Mishra, Viswanath Gopalakrishnan

机构 * International Institute of Information Technology Bangalore(国际信息科技学院班加罗尔)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

Comments First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21562 2025-08-05 cs.CL cs.AI cs.AR 76%

FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction

Jun Yin, Pengyu Zeng, Jing Zhong, Peilin Li, Miao Zhang, Ran Luo, Shuai Lu

专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19149 2025-06-12 cs.CV cs.AI cs.CR cs.LG 76%

Multimodal Pragmatic Jailbreak on Text-to-image Models

Tong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang, Shuo Chen, Philip Torr, Vera Demberg, Volker Tresp, Jindong Gu

机构 * LMU Munich, Germany(慕尼黑大学) Munich Center for Machine Learning, Germany(慕尼黑机器学习中心) Saarland University, Germany(萨尔兰大学) Max Planck Institute for Informatics, Germany(马克斯·普朗克信息研究所) Cornell University, USA(康奈尔大学) University of Oxford, UK(牛津大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08838 2025-05-20 eess.IV cs.AI cs.CV 76%

Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts

Peixuan Ge, Tongkun Su, Faqin Lv, Baoliang Zhao, Peng Zhang, Chi Hong Wong, Liang Yao, Yu Sun, Zenan Wang, Pak Kin Wong, Ying Hu

机构 * Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) University of Macau(澳门大学) Chinese PLA General Hospital(中国人民解放军总医院) Macau University of Science and Technology(澳门科学技术大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10453 2025-05-06 cs.CV cs.GR cs.MM 76%

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Liu He, Yizhi Song, Hejun Huang, Pinxin Liu, Yunlong Tang, Daniel Aliaga, Xin Zhou

机构 * Purdue University(普渡大学) Baidu USA(百度美国公司) University of Rochester(罗切斯特大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.MM

Comments Accepted by CVPR 2025 AI4CC Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04423 2025-04-08 cs.CV cs.AI 76%

UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Yang Jiao, Haibo Qiu, Zequn Jie, Shaoxiang Chen, Jingjing Chen, Lin Ma, Yu-Gang Jiang

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

Comments Accpeted to CVPR 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14492 2025-04-03 cs.CV cs.AI cs.LG cs.RO 76%

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

NVIDIA, :, Hassan Abu Alhaija, Jose Alvarez, Maciej Bala, Tiffany Cai, Tianshi Cao, Liz Cha, Joshua Chen, Mike Chen, Francesco Ferroni, Sanja Fidler, Dieter Fox, Yunhao Ge, Jinwei Gu, Ali Hassani, Michael Isaev, Pooya Jannaty, Shiyi Lan, Tobias Lasser, Huan Ling, Ming-Yu Liu, Xian Liu, Yifan Lu, Alice Luo, Qianli Ma, Hanzi Mao, Fabio Ramos, Xuanchi Ren, Tianchang Shen, Xinglong Sun, Shitao Tang, Ting-Chun Wang, Jay Wu, Jiashu Xu, Stella Xu, Kevin Xie, Yuchong Ye, Xiaodong Yang, Xiaohui Zeng, Yu Zeng

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16781 2025-04-01 cs.CV cs.AI 76%

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Yiheng Li, Ruibing Hou, Hong Chang, Shiguang Shan, Xilin Chen

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15217 2024-08-28 eess.IV cs.AI cs.CV 76%

Fundus2Video: Cross-Modal Angiography Video Generation from Static Fundus Photography with Clinical Knowledge Guidance

Weiyi Zhang, Siyu Huang, Jiancheng Yang, Ruoyu Chen, Zongyuan Ge, Yingfeng Zheng, Danli Shi, Mingguang He

专题命中 多模态生成 :cross-modal(title);分类 cs.CV、cs.AI

Comments The paper has been accepted by Medical Image Computing and Computer Assisted Intervention Society (MICCAI) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10005 2024-07-19 cs.CV cs.CL 76%

Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation

Kohei Uehara, Nabarun Goswami, Hanqin Wang, Toshiaki Baba, Kohtaro Tanaka, Tomohiro Hashimoto, Kai Wang, Rei Ito, Takagi Naoya, Ryo Umagami, Yingyi Wen, Tanachai Anakewat, Tatsuya Harada

专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10888 2024-07-16 eess.IV cs.AI cs.CV 76%

Leveraging Multimodal CycleGAN for the Generation of Anatomically Accurate Synthetic CT Scans from MRIs

Leonardo Crespi, Samuele Camnasio, Damiano Dei, Nicola Lambri, Pietro Mancosu, Marta Scorsetti, Daniele Loiacono

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

Comments Currently submitted to: Scientific Reports

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19589 2024-07-01 cs.SD cs.LG cs.MM eess.AS 76%

Network Bending of Diffusion Models for Audio-Visual Generation

Luke Dzwonczyk, Carmine Emanuele Cella, David Ban

专题命中 多模态生成 :audio-visual(title);分类 cs.MM、eess.AS

Comments 8 pages, 5 figures, to be published in the proceedings of the 27th International Conference on Digital Audio Effects (DAFx24), for additional image and video examples see https://dzluke.github.io/DAFX2024/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12410 2024-07-01 cs.CL cs.AI cs.LG q-bio.QM 76%

nach0: Multimodal Natural and Chemical Languages Foundation Model

Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov, Daniil Polykovskiy, Annika Brundyn, Aastha Jhunjhunwala, Anthony Costa, Alex Aliper, Alán Aspuru-Guzik, Alex Zhavoronkov

专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI

Comments Accepted to Chemical Science Journal. Models are publicly available via https://huggingface.co/insilicomedicine/nach0_base and https://huggingface.co/insilicomedicine/nach0_large

Journal ref Chemical Science, 15(22), 8380-8389, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19204 2024-05-01 cs.CV cs.AI cs.GR 76%

NeRF-Insert: 3D Local Editing with Multimodal Control Signals

Benet Oriol Sabat, Alessandro Achille, Matthew Trager, Stefano Soatto

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16806 2024-02-07 cs.CL cs.AI cs.LG 76%

Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking

David Bayani

专题命中 多模态生成 :cross-modal(title);分类 cs.CL、cs.AI

Comments Accepted in EACL 2024 as a long paper. See accepted/#long-papers" target="_blank" rel="noopener">https://2024.eacl.org/program/findings-accepted/#long-papers . Note: this paper's ArXiv version includes additional discussion, analysis, and types of experiments compared to the EACL version. Changes introduced in V2 of ArXiv paper: only this comment metadata. V1 was initially submission on July 26th, 2023 - release was delayed by ArXiv for a few days

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16882 2023-11-29 cs.CV cs.CL cs.LG 76%

Optimisation-Based Multi-Modal Semantic Image Editing

Bowen Li, Yongxin Yang, Steven McDonagh, Shifeng Zhang, Petru-Daniel Tudosiu, Sarah Parisot

专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07630 2023-11-15 cs.SD cs.CV cs.LG eess.AS 76%

Cross-modal Generative Model for Visual-Guided Binaural Stereo Generation

Zhaojian Li, Bin Zhao, Yuan Yuan

专题命中 多模态生成 :cross-modal(title);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00908 2022-11-30 cs.CL cs.CV 76%

Clue: Cross-modal Coherence Modeling for Caption Generation

Malihe Alikhani, Piyush Sharma, Shengjie Li, Radu Soricut, Matthew Stone

专题命中 多模态生成 :cross-modal(title);分类 cs.CV、cs.CL

Comments Accepted as a long paper to ACL 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12809 2022-06-22 cs.CV cs.AI eess.SP 76%

Warping of Radar Data into Camera Image for Cross-Modal Supervision in Automotive Applications

Christopher Grimm, Tai Fei, Ernst Warsitz, Ridha Farhoud, Tobias Breddermann, Reinhold Haeb-Umbach

专题命中 多模态生成 :cross-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01651 2021-12-06 cs.CV cs.AI 76%

Multi-modal application: Image Memes Generation

Zhiyuan Liu, Chuanzheng Sun, Yuxin Jiang, Shiqi Jiang, Mei Ming

专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.10980 2021-06-08 cs.CL cs.CV cs.LG stat.ML 76%

Multimodal Story Generation on Plural Images

Jing Jiang

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL

Comments This is an undergraduate project report. Completed Dec. 2019 at the Cooper Union

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09164 2021-01-19 cs.LG cs.AI cs.CV cs.RO 76%

Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders

Masha Itkina, Boris Ivanovic, Ransalu Senanayake, Mykel J. Kochenderfer, Marco Pavone

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

Comments 21 pages, 15 figures, 34th Conference on Neural Information Processing Systems (NeurIPS 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09522 2020-12-23 cs.CV cs.CL 76%

Multimodal Research in Vision and Language: A Review of Current and Emerging Trends

Shagun Uppal, Sarthak Bhagat, Devamanyu Hazarika, Navonil Majumdar, Soujanya Poria, Roger Zimmermann, Amir Zadeh

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.07385 2020-03-18 cs.CL cs.AI 76%

A Formal Analysis of Multimodal Referring Strategies Under Common Ground

Nikhil Krishnaswamy, James Pustejovsky

专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI

Comments 9 pages (incl refs), 7 figures, 3 tables, proceedings of LREC 2020 (postponed due to COVID-19)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.12352 2019-02-27 cs.CL cs.AI cs.LG cs.NE 76%

DialogWAE: Multimodal Response Generation with Conditional Wasserstein Auto-Encoder

Xiaodong Gu, Kyunghyun Cho, Jung-Woo Ha, Sunghun Kim

专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI

Comments Published as a conference paper at ICLR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13532 2024-10-02 eess.IV cs.CV 76%

Physics-Informed Latent Diffusion for Multimodal Brain MRI Synthesis

Sven Lüpke, Yousef Yeganeh, Ehsan Adeli, Nassir Navab, Azade Farshad

专题命中 多模态生成 :multimodal(title,comments);分类 cs.CV

Comments 5th International Workshop on Multiscale Multimodal Medical Imaging (MICCAI 2024), Project page: https://sven-luepke.github.io/phy-ldm-mri/

详情

展开后加载摘要…

URL PDF HTML 收藏