arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2507.17083 2025-07-24 cs.CV cs.AI 81%

SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction

Zaipeng Duan, Chenxu Dang, Xuzhong Hu, Pei An, Junfeng Ding, Jie Zhan, Yunbiao Xu, Jie Ma

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07611 2025-07-15 cs.CL cs.AI 81%

Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models

Shuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, Yida Xu, Yunya Song, Xian Yang

机构 * Hong Kong Baptist University(香港 Baptist 大学) Shanxi University(山西大学) Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海先进研究院) Hong Kong University of Science and Technology(香港科技大学) The University of Manchester(曼彻斯特大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 13 pages. 7 figures

Journal ref This paper is accpeted by ACL2025(Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08719 2025-07-14 cs.CL cs.AI cs.SE 81%

Multilingual Multimodal Software Developer for Code Generation

Linzheng Chai, Jian Yang, Shukai Liu, Wei Zhang, Liran Wang, Ke Jin, Tao Sun, Congnan Liu, Chenchen Zhang, Hualei Zhu, Jiaheng Liu, Xianjie Wu, Ge Zhang, Tianyu Liu, Zhoujun Li

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20214 2025-07-09 cs.CV cs.MM 81%

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

Yanzhe Chen, Huasong Zhong, Yan Li, Zhenheng Yang

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05626 2025-07-03 cs.CV cs.AI 81%

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models

Aarti Ghatkesar, Ganesh Venkatesh

机构 * AppliedML, Cerebras(应用机器学习,Cerebras)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18095 2025-06-24 cs.CV cs.AI cs.LG 81%

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Junying Chen, Zhenyang Cai, Pengcheng Chen, Shunian Chen, Ke Ji, Xidong Wang, Yunjin Yang, Benyou Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17218 2025-06-23 cs.CV cs.AI 81%

Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Zeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen, Chuang Gan

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Project page: https://vlm-mirage.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12325 2025-06-17 cs.SD cs.CL eess.AS 81%

GSDNet: Revisiting Incomplete Multimodal-Diffusion from Graph Spectrum Perspective for Conversation Emotion Recognition

Yuntao Shou, Jun Yao, Tao Meng, Wei Ai, Cen Chen, Keqin Li

机构 * College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南大学林业科技学院) Department of Computer Science, Anhui Normal University(计算机科学系,安徽师范大学) Future Technology Institute, South China University of Technology(未来技术研究院,华南理工大学) Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11380 2025-06-16 cs.CV cs.AI 81%

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation

Xiaoxin Lu, Ranran Haoran Zhang, Yusen Zhang, Rui Zhang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 18 pages, 10 figures; Accepted to ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07903 2025-06-16 cs.LG cs.AI cs.CV 81%

Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces

Kevin Rojas, Yuchen Zhu, Sichen Zhu, Felix X. -F. Ye, Molei Tao

机构 * Machine Learning Center, Georgia Institute of Technology, Atlanta, GA School of Mathematics, Georgia Institute of Technology, Atlanta, GA Department of Mathematics \& Statistics, SUNY Albany, NY

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ICML 2025. Code available at https://github.com/KevinRojas1499/Diffuse-Everything

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10328 2025-06-13 cs.CV cs.AI cs.LG 81%

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework

Sadia Kamal, Tim Oates, Joy Wan

机构 * Department of Computer Science, University of Maryland, Baltimore County(计算机科学系,马里兰大学巴尔的摩县分校) Department of Dermatology, Johns Hopkins University School of Medicine(皮肤科系,约翰霍普金斯大学医学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at IEEE/CVF Computer Society Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19593 2025-06-11 cs.CL cs.CV 81%

SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs

Xin Su, Man Luo, Kris W Pan, Tien Pei Chou, Vasudev Lal, Phillip Howard

机构 * Intel Labs(英特尔实验室) Amazon(亚马逊)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments ICML 2025 Spotlight Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07848 2025-06-10 cs.CV cs.AI 81%

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Teng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang, Yuan Zhou, Qinglin Lu, Ran Yi

机构 * Shanghai Jiao Tong University(上海交通大学) Tencent Hunyuan(腾讯文英) Zhejiang University(浙江大学)

专题命中 多模态生成 :cross-modal(title);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05896 2025-06-09 cs.RO cs.AI cs.CV 81%

Object Navigation with Structure-Semantic Reasoning-Based Multi-level Map and Multimodal Decision-Making LLM

Chongshang Yan, Jiaxuan He, Delun Li, Yi Yang, Wenjie Song

机构 * School of Automation(自动化学院) Beijing Institute of Technology(北京理工大学)

专题命中 多模态生成 :multimodal(title);MLLM(abstract);分类 cs.CV、cs.AI

Comments 16 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06542 2025-06-09 cs.CV cs.AI 81%

ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability

Jianwen Sun, Yukang Feng, Chuanhao Li, Fanrui Zhang, Zizhen Li, Jiaxin Ai, Sizhuo Zhou, Yu Dai, Shenglin Zhang, Kaipeng Zhang

机构 * Nankai University(南开大学) Shanghai Innovation Institute(上海创新研究院) University of Science and Technology of China(中国科学技术大学) Wuhan University(武汉大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03191 2025-06-05 cs.CV cs.AI 81%

Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward

Muhammad Islam, Tao Huang, Euijoon Ahn, Usman Naseem

机构 * College of Science and Engineering(科学与工程学院) James Cook University(詹姆斯库克大学) School of Computing(计算学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22222 2025-05-29 cs.CV cs.CL 81%

Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

Yunsoo Kim, Jinge Wu, Su-Hwan Kim, Pardeep Vasudev, Jiashu Shen, Honghan Wu

机构 * UCL(伦敦大学学院) Technical University of Munich(慕尼黑技术大学) University of Oxford(牛津大学) University of Glasgow(格拉斯哥大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21715 2025-05-29 eess.IV cs.AI cs.CV 81%

Privacy-Preserving Chest X-ray Report Generation via Multimodal Federated Learning with ViT and GPT-2

Md. Zahid Hossain, Mustofa Ahmed, Most. Sharmin Sultana Samu, Md. Rakibul Islam

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Preprint, manuscript under-review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18880 2025-05-27 cs.CV cs.AI 81%

REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing

Weihan Xu, Yimeng Ma, Jingyue Huang, Yang Li, Wenye Ma, Taylor Berg-Kirkpatrick, Julian McAuley, Paul Pu Liang, Hao-Wen Dong

机构 * Duke University(杜克大学) University of California, San Diego(加州大学圣迭戈分校) MBZUAI(马克斯·普朗克人工智能研究所) MIT(麻省理工学院) University of Michigan(密歇根大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18775 2025-05-27 cs.CV cs.AI 81%

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks

Jiayu Wang, Yang Jiao, Yue Yu, Tianwen Qian, Shaoxiang Chen, Jingjing Chen, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, Fudan University(上海智能信息处理关键实验室,复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17501 2025-05-26 cs.CV cs.AI cs.LG 81%

RoHyDR: Robust Hybrid Diffusion Recovery for Incomplete Multimodal Emotion Recognition

Yuehan Jin, Xiaoqing Liu, Yiyuan Yang, Zhiwen Yu, Tong Zhang, Kaixiang Yang

机构 * School of Computer Science and Engineering(计算机科学与工程学院) South China University of Technology(华南理工大学) School of Future Technology(未来技术学院) University of Oxford(牛津大学) Pengcheng Laboratory(鹏城实验室) Department of Computer Science(计算机科学系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17042 2025-05-26 cs.CL cs.CV cs.IR cs.LG 81%

VLM-KG: Multimodal Radiology Knowledge Graph Generation

Abdullah Abdullah, Seong Tae Kim

机构 * Department of Computer Science and Engineering, Kyung Hee University(计算机科学与工程系,庆熙大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12900 2025-05-20 cs.SE cs.AI cs.CG cs.CL cs.DB 81%

AutoGEEval: A Multimodal and Automated Framework for Geospatial Code Generation on GEE with Large Language Models

Shuyang Hou, Zhangxiao Shen, Huayi Wu, Jianyuan Liang, Haoyue Jiao, Yaxian Qing, Xiaopu Zhang, Xu Li, Zhipeng Gui, Xuefeng Guan, Longgang Xiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05415 2025-05-20 cs.CV cs.AI 81%

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding

Chenkai Xu, Xu Wang, Zhenyi Liao, Yishun Li, Tianqi Hou, Zhijie Deng

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei(华为) Tongji University(同济大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14066 2025-05-16 cs.CV cs.AI 81%

CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion

Nisha Huang, Weiming Dong, Yuxin Zhang, Fan Tang, Ronghui Li, Chongyang Ma, Xiu Li, Tong-Yee Lee, Changsheng Xu

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Pengcheng Laboratory(鹏城实验室) MAIS, Institute of Automation, Chinese Academy of Sciences(自动化所MAIS部) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) University of Chinese Academy of Sciences(中国科学院大学) ByteDance Inc.(字节跳动公司) National Cheng Kung University(国立成功大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01823 2025-05-14 cs.CV cs.AI cs.ET 81%

PhytoSynth: Leveraging Multi-modal Generative Models for Crop Disease Data Generation with Novel Benchmarking and Prompt Engineering Approach

Nitin Rai, Arnold W. Schumann, Nathan Boyd

机构 * Gulf Coast Research and Education Center (GCREC), University of Florida(佛罗里达大学 Gulf Coast Research and Education Center) Citrus Research and Education Center (CREC), University of Florida(佛罗里达大学 Citrus Research and Education Center)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01638 2025-05-06 eess.IV cs.AI cs.CV 81%

Seeing Heat with Color -- RGB-Only Wildfire Temperature Inference from SAM-Guided Multimodal Distillation using Radiometric Ground Truth

Michael Marinaccio, Fatemeh Afghah

机构 * Dept. of Electrical and Computer Engineering, Clemson University(电子与计算机工程系,克莱姆森大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 7 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19212 2025-04-29 cs.CV cs.AI 81%

CapsFake: A Multimodal Capsule Network for Detecting Instruction-Guided Deepfakes

Tuan Nguyen, Naseem Khan, Issa Khalil

机构 * Qatar Computing Research Institute(卡塔尔计算研究所) Hamad Bin Khalifa University(哈马德·本·卡尔法大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11707 2025-04-17 cs.CV cs.AI 81%

Towards Safe Synthetic Image Generation On the Web: A Multimodal Robust NSFW Defense and Million Scale Dataset

Muhammad Shahid Muneer, Simon S. Woo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Short Paper The Web Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11686 2025-04-17 cs.CV cs.AI 81%

Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics

Yiran He, Yun Cao, Bowen Yang, Zeyu Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 12 pages, 11 figures, 13IHMMSec2025

详情

展开后加载摘要…

URL PDF HTML 收藏