arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2510.05661 2025-10-08 cs.CV cs.MM 81%

When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach

Daniel Gonzálbez-Biosca, Josep Cabacas-Maso, Carles Ventura, Ismael Benito-Altamirano

机构 * eHealth Center, Faculty of Computer Science, Multimedia and Telecommunications, Universitat Oberta de Catalunya(eHealth中心,计算机科学、多媒体与电信学院,开放大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02403 2025-10-06 q-bio.QM cs.AI cs.CV 81%

Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model

Jalil Jalili, Yashraj Gavhane, Evan Walker, Anna Heinke, Christopher Bowd, Akram Belghith, Massimo A. Fazio, Christopher A. Girkin, C. Gustavo De Moraes, Jeffrey M. Liebmann, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17121 2025-10-06 cs.CL cs.AI 81%

NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation

Weiming Wu, Jin Ye, Zi-kang Wang, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo

机构 * School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学) National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学) School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26644 2025-10-01 cs.CV cs.AI cs.LG 81%

Stitch: Training-Free Position Control in Multimodal Diffusion Transformers

Jessica Bader, Mateusz Pach, Maria A. Bravo, Serge Belongie, Zeynep Akata

机构 * Technical University of Munich(慕尼黑技术大学) Helmholtz Munich(海德堡-慕尼黑研究所) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Copenhagen(哥本哈根大学)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21360 2025-09-29 cs.CV cs.AI 81%

Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models

Xingkai Peng, Jun Jiang, Meng Tong, Shuai Li, Weiming Zhang, Nenghai Yu, Kejiang Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21107 2025-09-26 cs.RO cs.AI cs.CV cs.LG 81%

Cross-Modal Instructions for Robot Motion Generation

William Barron, Xiaoxiang Dong, Matthew Johnson-Roberson, Weiming Zhi

机构 * College of Connected Computing, Vanderbilt University(连接计算学院,范德比尔特大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学) School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21476 2025-09-23 cs.CV cs.AI 81%

GarmentDiffusion: 3D Garment Sewing Pattern Generation with Multimodal Diffusion Transformers

Xinyu Li, Qi Yao, Yuanda Wang

机构 * Shenfu Research(沈孚研究所) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The 34th International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16197 2025-09-22 cs.CV cs.CL cs.LG 81%

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

Yanghao Li, Rui Qian, Bowen Pan, Haotian Zhang, Haoshuo Huang, Bowen Zhang, Jialing Tong, Haoxuan You, Xianzhi Du, Zhe Gan, Hyunjik Kim, Chao Jia, Zhenbang Wang, Yinfei Yang, Mingfei Gao, Zi-Yi Dou, Wenze Hu, Chang Gao, Dongxu Li, Philipp Dufter, Zirui Wang, Guoli Yin, Zhengdong Zhang, Chen Chen, Yang Zhao, Ruoming Pang, Zhifeng Chen

机构 * Apple(苹果公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15553 2025-09-22 cs.CV cs.AI stat.AP 81%

Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification

Tian Lan, Yiming Zheng, Jianxin Yin

机构 * School of Statistics, Renmin University of China(中国人民大学统计学院) Center for Applied Statistics and School of Statistics, Renmin University of China(中国人民大学应用统计中心和统计学院)

专题命中 多模态生成 :cross-modal(title);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01086 2025-09-18 cs.CV cs.AI 81%

DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing

Xiaolong Wang, Zhi-Qi Cheng, Jue Wang, Xiaojiang Peng

机构 * Shenzhen Technology University(深圳科技大学) University of Washington(华盛顿大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09070 2025-09-16 cs.LG cs.AI cs.CV 81%

FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院) Department of Knowledge Technologies, Jožef Stefan Institute(Jožef Stefan研究所知识技术系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08847 2025-09-12 cs.AI cs.CL cs.LG cs.SE 81%

Automated Unity Game Template Generation from GDDs via NLP and Multi-Modal LLMs

Amna Hassan

机构 * UET Taxila(塔希尔大学工程学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08489 2025-09-11 cs.CV cs.AI 81%

Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation

Kaleem Ahmad

机构 * Independent Researcher(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07817 2025-09-10 cs.CL cs.MM 81%

Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems

Xiaolin Chen, Xuemeng Song, Haokun Wen, Weili Guan, Xiangyu Zhao, Liqiang Nie

机构 * National University of Singapore(新加坡国立大学) Southern University of Science and Technology(南方科技大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) City University of Hong Kong(香港城市大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05263 2025-09-09 cs.AI cs.CV cs.LG 81%

LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation

Yinglin Duan, Zhengxia Zou, Tongwei Gu, Wei Jia, Zhan Zhao, Luyi Xu, Xinzhu Liu, Yenan Lin, Hao Jiang, Kang Chen, Shuang Qiu

机构 * NetEase, Inc.(网易公司) Beihang University(北京航空航天大学) Tsinghua University(清华大学) City University of Hong Kong(香港城市大学) Independent Researcher & Technical Artists(独立研究者及技术艺术家)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05714 2025-09-09 cs.AI cs.CV 81%

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs

Zhaoyu Fan, Kaihang Pan, Mingze Zhou, Bosheng Qin, Juncheng Li, Shengyu Zhang, Wenqiao Zhang, Siliang Tang, Fei Wu, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03535 2025-09-05 cs.CL cs.AI 81%

QuesGenie: Intelligent Multimodal Question Generation

Ahmed Mubarak, Amna Ahmed, Amira Nasser, Aya Mohamed, Fares El-Sadek, Mohammed Ahmed, Ahmed Salah, Youssef Sobhy

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 8 figures, 12 tables. Supervised by Dr. Ahmed Salah and TA Youssef Sobhy

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19320 2025-08-29 cs.CV cs.AI 81%

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Ming Chen, Liyuan Cui, Wenyuan Zhang, Haoxian Zhang, Yan Zhou, Xiaohan Li, Songlin Tang, Jiwen Liu, Borui Liao, Hejia Chen, Xiaoqiang Liu, Pengfei Wan

机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队) Zhejiang University(浙江大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Technical Report. Project Page: https://chenmingthu.github.io/milm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15194 2025-08-27 cs.CV cs.AI cs.LG 81%

DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models

Sungnyun Kim, Junsoo Lee, Kibeom Hong, Daesik Kim, Namhyuk Ahn

机构 * KAIST AI(韩国科学技术院人工智能研究所) NAVER WEBTOON AI Sookmyung Women’s University(成均馆女子大学) Inha University(釜山大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Expert Systems with Applications 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16930 2025-08-26 eess.AS cs.CV cs.SD 81%

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Sizhe Shan, Qiulin Li, Yutao Cui, Miles Yang, Yuehai Wang, Qun Yang, Jin Zhou, Zhao Zhong

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03001 2025-08-26 cs.CV cs.MM 81%

One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning

Hao Sun, Yu Song, Jiaqing Liu, Jihong Hu, Yen-Wei Chen, Lanfen Lin

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Information Science and Engineering, Ritsumeikan University(立命馆大学信息科学与工程学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05658 2025-08-12 cs.CR cs.CV cs.MM 81%

Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards

Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang, Zhengyu Zhao, Zheng Wang

机构 * Information Engineering University Zhengzhou China School of Computer Science, \ University Wuhan China Xi’an Jiaotong University Xi’an China Wuhan University Wuhan China Information Engineering University School of Computer Science, \ University Xi’an Jiaotong University Wuhan University

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06492 2025-08-11 cs.CV cs.CL 81%

Effective Training Data Synthesis for Improving MLLM Chart Understanding

Yuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li, Gaowen Liu, Ali Payani, Yuan-Sen Ting, Liang Zheng

机构 * Australian National University(澳大利亚国立大学) Ohio State University(俄亥俄州立大学) Cisco(思科公司) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICCV 2025 (poster). 26 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06510 2025-08-08 cs.CV cs.AI 81%

AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis

Shidan He, Lei Liu, Xiujun Shu, Bo Wang, Yuanhao Feng, Shen Zhao

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03069 2025-08-08 cs.CV cs.AI 81%

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Liao Qu, Huichao Zhang, Yiheng Liu, Xu Wang, Yi Jiang, Yiming Gao, Hu Ye, Daniel K. Du, Zehuan Yuan, Xinglong Wu

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR 2025; Code and models: https://github.com/ByteVisionLab/TokenFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21167 2025-08-07 cs.CV cs.AI 81%

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions

Donglu Yang, Liang Zhang, Zihao Yue, Liangyu Chen, Yichen Xu, Wenxuan Wang, Qin Jin

机构 * independent researcher(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05683 2025-08-07 cs.LG cs.AI cs.CR cs.MM 81%

Multi-Modal Multi-Task Federated Foundation Models for Next-Generation Extended Reality Systems: Towards Privacy-Preserving Distributed Intelligence in AR/VR/MR

Fardis Nadimi, Payam Abdisarabshali, Kasra Borazjani, Jacob Chakareski, Seyyedali Hosseinalipour

机构 * University at Buffalo–SUNY(布法罗大学-纽约州立大学) Department of Electrical Engineering(电气工程系) New Jersey Institute of Technology (NJIT)(新泽西理工学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments 16 pages, 4 Figures, 8 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03426 2025-08-06 cs.CV cs.AI cs.LG 81%

R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation

Futian Wang, Yuhan Qiao, Xiao Wang, Fuling Wang, Yuxiang Zhang, Dengdi Sun

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23058 2025-08-01 cs.CV cs.AI 81%

Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

Alexandru Buburuzan

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments A dissertation submitted to The University of Manchester for the degree of Bachelor of Science in Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22920 2025-08-01 cs.CL cs.AI 81%

Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey

Jindong Li, Yali Fu, Jiahong Liu, Linxiao Cao, Wei Ji, Menglin Yang, Irwin King, Ming-Hsuan Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏