arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2508.02317 2025-08-08 cs.CL cs.AI cs.DC 62%

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Qianli Ma, Yaowei Zheng, Zhelun Shi, Zhongkai Zhao, Bin Jia, Ziyue Huang, Zhiqi Lin, Youjie Li, Jiacheng Yang, Yanghua Peng, Zhi Zhang, Xin Liu

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :omni-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16978 2025-08-08 cs.CV cs.AI 62%

PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask

Jeongho Kim, Hoiyeong Jin, Sunghyun Park, Jaegul Choo

机构 * KAIST(韩国科学技术院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03696 2025-08-07 cs.CR cs.AI cs.CV 62%

PLA: Prompt Learning Attack against Text-to-Image Generative Models

Xinqi Lyu, Yihao Liu, Yanjie Li, Bin Xiao

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 3 figures, and published to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02362 2025-08-05 cs.CV cs.AI 62%

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

Xu Wang, Shengeng Tang, Fei Wang, Lechao Cheng, Dan Guo, Feng Xue, Richang Hong

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20987 2025-07-30 cs.CV cs.AI 62%

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1

Xinhan Di, Kristin Qi, Pengqian Yu

机构 * Computer Science, University of Massachusetts Boston(马萨诸塞大学波士顿分校计算机科学系) National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments WiCV @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20536 2025-07-30 cs.CV cs.AI cs.HC 62%

T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation

Chieh-Yun Chen, Min Shi, Gong Zhang, Humphrey Shi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19492 2025-07-29 cs.HC cs.AI cs.CV 62%

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation

Jovana Kondic, Pengyuan Li, Dhiraj Joshi, Zexue He, Shafiq Abedin, Jennifer Sun, Ben Wiesel, Eli Schwartz, Ahmed Nassar, Bo Wu, Assaf Arbelle, Aude Oliva, Dan Gutfreund, Leonid Karlinsky, Rogerio Feris

机构 * MIT(麻省理工学院) MIT-IBM Watson AI Labs(麻省理工-IBM沃森人工智能实验室) IBM Research(IBM研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00045 2025-07-29 cs.MM cs.AI cs.LG 62%

Detecting Multimedia Generated by Large AI Models: A Survey

Li Lin, Neeraj Gupta, Yue Zhang, Hainan Ren, Chun-Hao Liu, Feng Ding, Xin Wang, Xin Li, Luisa Verdoliva, Shu Hu

机构 * Department of Computer and Information Technology, Purdue University(普渡大学计算机与信息科技系) School of Software, Nanchang University(南昌大学软件学院) Amazon Prime Video(亚马逊Prime视频) Department of Epidemiology and Biostatistics, School of Public Health(公共卫生学院流行病学与生物统计学系) Department of Computer Science, College of Nanotechnology, Science, and Engineering(纳米技术、科学与工程学院计算机科学系) University at Albany, SUNY(阿尔巴尼大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18667 2025-07-28 cs.CV cs.AI 62%

Gen-AI Police Sketches with Stable Diffusion

Nicholas Fidalgo, Aaron Contreras, Katherine Harvey, Johnny Ni

机构 * Harvard College(哈佛学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14575 2025-07-22 cs.CV cs.AI 62%

Benchmarking GANs, Diffusion Models, and Flow Matching for T1w-to-T2w MRI Translation

Andrea Moschetto, Lemuel Puglisi, Alec Sargood, Pierluigi Dell'Acqua, Francesco Guarnera, Sebastiano Battiato, Daniele Ravì

机构 * Università degli Studi di Catania(卡塔尼亚大学) Università degli Studi di Messina(梅斯纳大学) University College London(伦敦大学学院)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17562 2025-07-21 cs.CV cs.CL 62%

LLM-driven Medical Report Generation via Communication-efficient Heterogeneous Federated Learning

Haoxuan Che, Haibo Jin, Zhengrui Guo, Yi Lin, Cheng Jin, Hao Chen

机构 * Department of Computer Science and Engineering, Department of Chemical and Biological Engineering and Division of Life Science, Hong Kong University of Science and Technology(计算机科学与工程系、化学与生物工程系及生命科学分校,香港科学与技术大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by IEEE TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12761 2025-07-18 cs.CV cs.AI 62%

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

Hanlei Shi, Leyuan Qu, Yu Liu, Di Gao, Yuhua Zheng, Taihao Li

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(杭州高等研究院,中国科学院大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11694 2025-07-17 cs.CL cs.AI 62%

ExpliCIT-QA: Explainable Code-Based Image Table Question Answering

Maximiliano Hormazábal Lagos, Álvaro Bueno Sáez, Pedro Alonso Doval, Jorge Alcalde Vesteiro, Héctor Cerezo-Costas

机构 * Computer Vision Center, Universitat Autònoma de Barcelona(计算机视觉中心,巴塞罗那自治大学) Gradiant

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments This work has been accepted for presentation at the 24nd Portuguese Conference on Artificial Intelligence (EPIA 2025) and will be published in the proceedings by Springer in the Lecture Notes in Computer Science (LNCS) series. Please cite the published version when available

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11152 2025-07-16 eess.IV cs.AI cs.CV 62%

Latent Space Consistency for Sparse-View CT Reconstruction

Duoyou Chen, Yunqing Chen, Can Zhang, Zhou Wang, Cheng Chen, Ruoxiu Xiao

机构 * University of Science and Technology Beijing(北京科技大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ACMMM2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11015 2025-07-16 cs.CV cs.AI 62%

Semantically Informed Salient Regions Guided Radiology Report Generation

Zeyi Hou, Zeqiang Wei, Ruixin Yan, Ning Lang, Xiuzhuang Zhou

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学) Department of Radiology, Peking University Third Hospital(北京大学第三医院放射科)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24085 2025-07-15 cs.CV cs.AI 62%

Imagine for Me: Creative Conceptual Blending of Real Images and Text via Blended Attention

Wonwoong Cho, Yanxia Zhang, Yan-Ying Chen, David I. Inouye

机构 * Elmore Family School of Electrical and Computer Engineering, Purdue University(埃尔摩家庭电气与计算机工程学院,普渡大学) Toyota Research Institute(丰田研究院)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Project website is available at https://imagineforme.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19021 2025-07-15 cs.CV cs.AI 62%

Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

Tao Liu, Rongjie Li, Chongyu Wang, Xuming He

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08039 2025-07-14 cs.CV cs.AI cs.LG 62%

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models

Sujith Vemishetty, Advitiya Arora, Anupama Sharma

机构 * Synechron

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05621 2025-07-09 cs.CV cs.MM 62%

AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework

Suoxiang Zhang, Xiaxi Li, Hongrui Chang, Zhuoyan Hou, Guoxin Wu, Ronghua Ji

机构 * China Agricultural University(中国农业大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05302 2025-07-09 cs.CV cs.AI 62%

CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection

Binjia Zhou, Hengrui Lou, Lizhe Chen, Haoyuan Li, Dawei Luo, Shuai Chen, Jie Lei, Zunlei Feng, Yijun Bei

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) School of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Ant Group(蚂蚁集团) School of Computer Science and Technology, Zhejiang University of Technology(浙江工业大学计算机科学与技术学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19179 2025-07-08 cs.CV cs.AI cs.LG 62%

Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning

Dongwei Sun, Jing Yao, Wu Xue, Changsheng Zhou, Pedram Ghamisi, Xiangyong Cao

机构 * School of Computer Science and Technology and the Ministry of Education Key Lab for Intelligent Networks and Network Security, Xi’an Jiaotong University(计算机科学与技术学院和教育部智能网络与网络安全重点实验室,西安交通大学) Aerospace Information Research Institute, Chinese Academy of Sciences(航天信息研究所,中国科学院) Space Engineering University(航天工程大学) School of Mathematics and Statistics, Guangdong University of Technology(数学与统计学院,广东工业大学) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍兹中心)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04522 2025-07-08 cs.CV cs.AI cs.RO 62%

Grounded Gesture Generation: Language, Motion, and Space

Anna Deichler, Jim O'Regan, Teo Guichoux, David Johansson, Jonas Beskow

机构 * KTH Royal Institute of Technology(皇家理工学院) Sorbonne University(索邦大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted as a non-archival paper at the CVPR 2025 Humanoid Agents Workshop. Project page: https://groundedgestures.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04954 2025-07-08 cs.CV cs.CL cs.LG 62%

Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation

Xi Zhang, Zaiqiao Meng, Jake Lever, Edmond S. L. Ho

机构 * School of Computing Science, University of Glasgow(计算科学学院,格拉斯哥大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by BioNLP@ACL 2024

Journal ref Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, ACL 2024, pp. 624-634

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03313 2025-07-08 cs.CV cs.AI 62%

Personalized Image Generation from an Author Writing Style

Sagar Gandhi, Vishal Gandhi

机构 * Joyspace AI

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01055 2025-07-03 eess.IV cs.AI cs.CV 62%

Prompt Mechanisms in Medical Imaging: A Comprehensive Survey

Hao Yang, Xinlong Liang, Zhang Li, Yue Sun, Zheyu Hu, Xinghe Xie, Behdad Dashtbozorg, Jincheng Huang, Shiwei Zhu, Luyi Han, Jiong Zhang, Shanshan Wang, Ritse Mann, Qifeng Yu, Tao Tan

机构 * Faculty of Applied Sciences, Macao Polytechnic University(应用科学学院,澳门理工学院) Netherlands Cancer Institute, Department of Radiology(荷兰癌症研究所,放射科) Radboud University Medical Centre, Department of Radiology and Nuclear Medicine(拉德堡德大学医学中心,放射科和核医学科) College of Aerospace Science and Engineering, National University of Defense Technology(航空航天科学与工程学院,国防科技大学) Medical Department of Breast Cancer, Hunan Cancer Hospital(乳腺癌医学部,湖南癌症医院) the Affiliated Cancer Hospital of Xiangya School of Medicine, Central South University(湘雅医学院附属肿瘤医院,中南大学) Faculty of Biomedical Engineering, Eindhoven University of Technology(生物医学工程学院,埃因霍温理工大学) Laboratory of Advanced Theranostic Materials and Technology, University of Chinese Academy of Sciences(先进诊疗材料与技术实验室,中国科学院大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10493 2025-07-01 cs.CV cs.AI cs.LG 62%

AlignGuard: Scalable Safety Alignment for Text-to-Image Generation

Runtao Liu, I Chieh Chen, Jindong Gu, Jipeng Zhang, Renjie Pi, Qifeng Chen, Philip Torr, Ashkan Khakzar, Fabio Pizzati

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23641 2025-07-01 cs.CV cs.AI 62%

VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation

Peng Huang, Junhu Fu, Bowen Guo, Zeju Li, Yuanyuan Wang, Yi Guo

机构 * College of Biomedical Engineering, Fudan University, Shanghai 200433, China(复旦大学生物医学工程学院) Key Laboratory of Medical Imaging Computing and Computer Assisted Intervention of Shanghai, Shanghai 200032, China(上海医学影像计算与计算机辅助干预重点实验室)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21655 2025-06-30 cs.LG cs.AI cs.CV 62%

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Minjie Hong, Zirun Guo, Yan Xia, Zehan Wang, Ziang Zhang, Tao Jin, Zhou Zhao

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18309 2025-06-27 cs.CV cs.AI 62%

MvKeTR: Chest CT Report Generation with Multi-View Perception and Knowledge Enhancement

Xiwei Deng, Xianchun He, Jianfeng Bao, Yudan Zhou, Shuhui Cai, Congbo Cai, Zhong Chen

机构 * Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院) Department of Magnetic Resonance Imaging, The First Affiliated Hospital of Zhengzhou University(郑州大学第一附属医院磁共振成像科) Department of Electronic Science, Xiamen University(厦门大学电子科学系)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in IEEE Journal of Biomedical and Health Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20178 2025-06-26 cs.CL cs.AI cs.LG 62%

COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees

Zhiyuan Wang, Jinhao Duan, Qingni Wang, Xiaofeng Zhu, Tianlong Chen, Xiaoshuang Shi, Kaidi Xu

机构 * University of Electronic Science and Technology of China(电子科技大学) Drexel University(德雷塞尔大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏