arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2501.13925 2025-01-24 cs.CV 79%

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad S. Khan, Salman Khan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12173 2025-01-22 cs.CV 79%

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

Shiyue Zhang, Zheng Chong, Xi Lu, Wenqing Zhang, Haoxiang Li, Xujie Zhang, Jiehui Huang, Xiao Dong, Xiaodan Liang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11233 2025-01-22 cs.IR cs.CL cs.MA 79%

PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents

Kanika Goswami, Puneet Mathur, Ryan Rossi, Franck Dernoncourt

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at ECIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07086 2025-01-14 cs.CL 79%

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models

Yongyu Mu, Hengyu Li, Junxin Wang, Xiaoxuan Zhou, Chenglong Wang, Yingfeng Luo, Qiaozhi He, Tong Xiao, Guocheng Chen, Jingbo Zhu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02167 2025-01-07 cs.CV 79%

Generating Multimodal Images with GAN: Integrating Text, Image, and Style

Chaoyi Tan, Wenqing Zhang, Zhen Qi, Kowei Shih, Xinshi Li, Ao Xiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20725 2024-12-31 cs.CV 79%

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

Min Zhang, Zilin Wang, Liyan Chen, Kunhong Liu, Juncong Lin

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08603 2024-12-31 cs.CL 79%

A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

Hanguang Xiao, Feizhong Zhou, Xingyue Liu, Tianqi Liu, Zhipeng Li, Xin Liu, Xiaoxuan Huang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Journal ref Information Fusion, 117 (2025) 102888

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18928 2024-12-30 cs.CV cs.LG 79%

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation

Lunhao Duan, Shanshan Zhao, Wenjun Yan, Yinglun Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Mingming Gong, Gui-Song Xia

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17225 2024-12-24 cs.CV 79%

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Lichen Ma, Tiezhu Yue, Pei Fu, Yujie Zhong, Kai Zhou, Xiaoming Wei, Jie Hu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13129 2024-12-24 cs.CV cs.LG 79%

M3T: Multi-Modal Medical Transformer to bridge Clinical Context with Visual Insights for Retinal Image Medical Description Generation

Nagur Shareef Shaik, Teja Krishna Cherukuri, Dong Hye Ye

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments This paper has been accepted for presentation at the IEEE International Conference on Image Processing (ICIP 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04363 2024-12-19 cs.CV 79%

Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs

Junhao Chen, Xiang Li, Xiaojun Ye, Chao Li, Zhaoxin Fan, Hao Zhao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by COLING 2025 (The 31st International Conference on Computational Linguistics) Project Page: https://idea23d.github.io/ Code: https://github.com/yisuanwang/Idea23D

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11849 2024-12-17 eess.IV cs.CV cs.LG 79%

Ensemble Learning and 3D Pix2Pix for Comprehensive Brain Tumor Analysis in Multimodal MRI

Ramy A. Zeineldin, Franziska Mathis-Ullrich

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted at the MICCAI BraTS Challenge 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01859 2024-12-17 eess.IV cs.CV 79%

A Novel Deep Learning Tractography Fiber Clustering Framework for Functionally Consistent White Matter Parcellation Using Multimodal Diffusion MRI and Functional MRI

Jin Wang, Bocheng Guo, Yijie Li, Junyi Wang, Yuqian Chen, Jarrett Rushmore, Nikos Makris, Yogesh Rathi, Lauren J O'Donnell, Fan Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12154 2024-12-17 cs.CV 79%

StyleBooth: Image Style Editing with Multimodal Instruction

Zhen Han, Chaojie Mao, Zeyinzi Jiang, Yulin Pan, Jingfeng Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02160 2024-12-13 eess.IV cs.CV 79%

MRI to PET Cross-Modality Translation using Globally and Locally Aware GAN (GLA-GAN) for Multi-Modal Diagnosis of Alzheimer's Disease

Apoorva Sikka, Skand Peri, Jitender Singh Virk, Usma Niyaz, Deepti R. Bathula

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07141 2024-12-11 cs.CV 79%

Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation

Qianhao Han, Junyi Liu, Zengchang Qin, Zheng Zheng

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted in IEEE Big Data 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06840 2024-12-11 cs.LG cs.CV 79%

MDiFF: Exploiting Multimodal Score-based Diffusion Models for New Fashion Product Performance Forecasting

Andrea Avogaro, Luigi Capogrosso, Franco Fummi, Marco Cristani

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at the FashionAI workshop @ the European Conference on Computer Vision (ECCV) 2024. arXiv admin note: substantial text overlap with arXiv:2412.05566

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05605 2024-12-10 cs.CV 79%

RefSAM3D: Adapting SAM with Cross-modal Reference for 3D Medical Image Segmentation

Xiang Gao, Kai Lu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05566 2024-12-10 cs.CV cs.LG 79%

Dif4FF: Leveraging Multimodal Diffusion Models and Graph Neural Networks for Accurate New Fashion Product Performance Forecasting

Andrea Avogaro, Luigi Capogrosso, Franco Fummi, Marco Cristani

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at the 27th International Conference on Pattern Recognition (ICPR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00499 2024-12-09 cs.CV cs.ET cs.LG eess.SP 79%

Cross-modal semantic segmentation for indoor environmental perception using single-chip millimeter-wave radar raw data

Hairuo Hu, Haiyong Cong, Zhuyu Shao, Yubo Bi, Jinghao Liu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments 5291 words, 17 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03809 2024-12-06 cs.CV 79%

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM

Quang Nguyen, Truong Vu, Trong-Tung Nguyen, Yuxin Wen, Preston K Robinette, Taylor T Johnson, Tom Goldstein, Anh Tran, Khoi Nguyen

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01407 2024-12-04 cs.CV 79%

HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving

Zehuan Wu, Jingcheng Ni, Xiaodong Wang, Yuxin Guo, Rui Chen, Lewei Lu, Jifeng Dai, Yuwen Xiong

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00243 2024-12-03 cs.RO cs.AI 79%

Realistic Corner Case Generation for Autonomous Vehicles with Multimodal Large Language Model

Qiujing Lu, Meng Ma, Ximiao Dai, Xuanhan Wang, Shuo Feng

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18850 2024-12-02 cs.CV 79%

CrossTracker: Robust Multi-modal 3D Multi-Object Tracking via Cross Correction

Lipeng Gu, Xuefeng Yan, Weiming Wang, Honghua Chen, Dingkun Zhu, Liangliang Nan, Mingqiang Wei

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00448 2024-11-21 cs.CV 79%

MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation

Xujie Zhang, Ente Lin, Xiu Li, Yuxuan Luo, Michael Kampffmeyer, Xin Dong, Xiaodan Liang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16273 2024-11-05 cs.CV 79%

M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

Mingshuang Luo, Ruibing Hou, Zhuo Li, Hong Chang, Zimo Liu, Yaowei Wang, Shiguang Shan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at NeurIPS 2024, 21 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23905 2024-11-01 cs.CV 79%

Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model

Hao Zhang, Lei Cao, Jiayi Ma

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by the 38th Conference on Neural Information Processing Systems (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12647 2024-10-28 cs.CV cs.RO 79%

DiffusionNOCS: Managing Symmetry and Uncertainty in Sim2Real Multi-Modal Category-level Pose Estimation

Takuya Ikeda, Sergey Zakharov, Tianyi Ko, Muhammad Zubair Irshad, Robert Lee, Katherine Liu, Rares Ambrus, Koichi Nishiwaki

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages. 9 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16942 2024-10-24 eess.IV cs.CV 79%

PASTA: Pathology-Aware MRI to PET Cross-Modal Translation with Diffusion Models

Yitong Li, Igor Yakushev, Dennis M. Hedderich, Christian Wachinger

专题命中 多模态生成 :cross-modal(title);multi-modal(abstract);分类 cs.CV

Journal ref Medical Image Computing and Computer Assisted Intervention (MICCAI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17032 2024-10-23 cs.AI 79%

Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups

Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel, Zoe Ashwood, Aida Mostafazadeh Davani, Mark Diaz, Michela Paganini, Alicia Parrish, Ding Wang, Vinodkumar Prabhakaran, Lora Aroyo, Verena Rieser

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏