arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2510.00046 2025-10-02 cs.CV cs.AI 62%

Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models

Xiaotian Zou

机构 * Xiaotian Zou

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25817 2025-10-01 cs.CL cs.CV 62%

Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer

Jaeyoung Kim, Jongho Lee, Hongjun Choi, Sion Jang

机构 * Teamreboott Inc.(Teamreboott公司) MIRI D.I.H Inc.(MIRI D.I.H公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22940 2025-09-30 cs.CL cs.CV 62%

LLMs Behind the Scenes: Enabling Narrative Scene Illustration

Melissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun, Alex Calderwood, Max Kreminski

机构 * Midjourney Northwestern University(西北大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21887 2025-09-29 cs.CV cs.MM 62%

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

Liyang Chen, Tianze Zhou, Xu He, Boshi Tang, Zhiyong Wu, Yang Huang, Yang Wu, Zhongqian Sun, Wei Yang, Helen Meng

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21375 2025-09-29 cs.CV cs.AI 62%

Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis

Aleksa Jelaca, Ying Jiao, Chang Tian, Marie-Francine Moens

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments text-to-image generation, automatic prompt, DPO, Counterfactual

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16663 2025-09-26 cs.CL cs.AI 62%

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

Yujin Han, Hao Chen, Andi Han, Zhiheng Wang, Xinyu Liu, Yingya Zhang, Shiwei Zhang, Difan Zou

专题命中 多模态生成 :MLLM(abstract);分类 cs.CL、cs.AI

Comments 31 pages, 16 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18907 2025-09-26 cs.AI cs.CV cs.RO 62%

EC-Diffuser: Multi-Object Manipulation via Entity-Centric Behavior Generation

Carl Qi, Dan Haramati, Tal Daniel, Aviv Tamar, Amy Zhang

机构 * UT Austin(得克萨斯大学) Technion, Israel Institute of Technology(技术学院,以色列技术学院) Brown University(布朗大学) Meta AI

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18179 2025-09-24 cs.CV cs.AI 62%

The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes

Sai Varun Kodathala, Rakesh Vunnam

机构 * Sports Vision, Inc.(体育视觉公司) Vizworld, Inc.(Vizworld公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13794 2025-09-22 cs.CV cs.AI 62%

LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation

Yang Zhou, Shiyu Zhao, Yuxiao Chen, Zhenting Wang, Can Jin, Dimitris N. Metaxas

机构 * Rutgers University(新泽西罗格斯大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15270 2025-09-22 cs.CV cs.AI 62%

PRISM: Phase-enhanced Radial-based Image Signature Mapping framework for fingerprinting AI-generated images

Emanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci, Roberto Di Pietro

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12888 2025-09-17 cs.CV cs.AI 62%

Runge-Kutta Approximation and Decoupled Attention for Rectified Flow Inversion and Semantic Editing

Weiming Chen, Zhihan Zhu, Yijia Wang, Zhihai He

机构 * Southern University of Science and Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12534 2025-09-17 eess.IV cs.AI cs.CV 62%

DeepEyeNet: Generating Medical Report for Retinal Images

Jia-Hong Huang

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments The paper is accepted by the Conference on Information and Knowledge Management (CIKM), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11661 2025-09-16 cs.CV cs.AI 62%

DTGen: Generative Diffusion-Based Few-Shot Data Augmentation for Fine-Grained Dirty Tableware Recognition

Lifei Hao, Yue Cheng, Baoqi Huang, Bing Jia, Xuandong Zhao

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10845 2025-09-16 cs.CL cs.MM 62%

Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production

Liqian Feng, Lintao Wang, Kun Hu, Dehui Kong, Zhiyong Wang

机构 * School of Computer Science(计算机科学学院) School of Science(科学学院) Faculty of Information Technology(信息技术学院)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10345 2025-09-16 cs.CV cs.AI 62%

Towards Understanding Visual Grounding in Visual Language Models

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所) Heriot-Watt University(赫瑞-沃德大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18985 2025-09-16 cs.LG cs.CL cs.CV 62%

STRICT: Stress Test of Rendering Images Containing Text

Tianyu Zhang, Xinyu Wang, Lu Li, Zhenghan Tai, Jijun Chi, Jingrui Tian, Hailin He, Suyuchen Wang

机构 * Mila, University of Montreal(蒙特利尔大学Mila) McGill University(麦吉尔大学) University of Pennsylvania(宾夕法尼亚大学) University of Toronto(多伦多大学) University of California, Los Angeles(加州大学洛杉矶分校) Southwestern University of Finance and Economics(西南财经大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08442 2025-09-11 cs.CV cs.AI cs.LG q-bio.NC 62%

Spherical Brownian Bridge Diffusion Models for Conditional Cortical Thickness Forecasting

Ivan Stoyanov, Fabian Bongratz, Christian Wachinger

机构 * Lab for AI in Medical Imaging, Technical University of Munich, Munich, Germany(人工智能医学影像实验室,慕尼黑技术大学,慕尼黑,德国) Munich Center for Machine Learning, Munich, Germany(慕尼黑机器学习中心,慕尼黑,德国)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05333 2025-09-09 cs.CV cs.AI 62%

RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness

Junghyun Park, Tuan Anh Nguyen, Dugki Min

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05321 2025-09-09 cs.CV cs.AI 62%

A Dataset Generation Scheme Based on Video2EEG-SPGN-Diffusion for SEED-VD

Yunfei Guo, Tao Zhang, Wu Huang, Yao Song

机构 * Chengdu Techman Software Co., Ltd.(成都技漫软件有限公司) Sichuan University(四川大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02278 2025-09-03 cs.GR cs.AI cs.MM 62%

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

Zikai Huang, Yihan Zhou, Xuemiao Xu, Cheng Xu, Xiaofen Xing, Jing Qin, Shengfeng He

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Centre for Smart Health, Hong Kong Polytechnic University(香港理工大学智能健康研究中心) CAS-Hong Kong Joint Laboratory for Multimodal Medical Molecular Imaging(中国科学院-香港联合多模态医学分子成像联合实验室) School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息学院) Pazhou Lab, Guangzhou, Guangdong, China(广州琶洲实验室) School of Computing and Information Systems, Singapore Management University(新加坡国立大学计算机与信息系统学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01656 2025-09-03 cs.CV cs.CL 62%

Reinforced Visual Perception with Tools

Zetong Zhou, Dongping Chen, Zixian Ma, Zhihan Hu, Mingyang Fu, Sinan Wang, Yao Wan, Zhou Zhao, Ranjay Krishna

机构 * ONE Lab, HUST(华中科技大学 ONE 实验室) ONE Lab, HUST University of Maryland(华中科技大学 与 马里兰大学 ONE 实验室) University of Washington(华盛顿大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17342 2025-08-26 cs.GR cs.CV cs.MM cs.SD 62%

DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions

Hengyuan Zhang, Zhe Li, Xingqun Qi, Mengze Li, Muyi Sun, Man Zhang, Sirui Han

机构 * Peking University(北京大学) The Hong Kong University of Science and Technology(香港科技大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.MM

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17062 2025-08-26 cs.CV cs.AI 62%

SSG-Dit: A Spatial Signal Guided Framework for Controllable Video Generation

Peng Hu, Yu Gu, Liang Luo, Fuji Ren

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(计算机科学与工程学院,电子科技大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13307 2025-08-15 cs.CV cs.AI 62%

Quantitative Comparison of Fine-Tuning Techniques for Pretrained Latent Diffusion Models in the Generation of Unseen SAR Images

Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Olivier Lévêque, Elise Colin

机构 * Paris-Saclay University(巴黎-萨克雷大学) ONERA - The French Aerospace Lab(法国航空航天实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08601 2025-08-15 cs.CV cs.AI 62%

Yan: Foundational Interactive Video Generation

Deheng Ye, Fangyun Zhou, Jiacheng Lv, Jianqi Ma, Jun Zhang, Junyan Lv, Junyou Li, Minwen Deng, Mingyu Yang, Qiang Fu, Wei Yang, Wenkai Lv, Yangbin Yu, Yewen Wang, Yonghang Guan, Zhihao Hu, Zhongbin Fang, Zhongqian Sun

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05296 2025-08-14 cs.AI cs.HC cs.SD eess.AS 62%

Revisiting Your Memory: Reconstruction of Affect-Contextualized Memory via EEG-guided Audiovisual Generation

Joonwoo Kwon, Heehwan Wang, Jinwoo Lee, Sooyoung Kim, Shinjae Yoo, Yuewei Lin, Jiook Cha

机构 * Michigan State University(密歇根州立大学) Seoul National University(首尔国立大学) Rutgers University(罗格斯大学) Brookhaven National Laboratory(布鲁赫斯国家实验室)

专题命中 多模态生成 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments Accepted at the ACM MM 2025 - The 1st CogMAEC Workshop (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09461 2025-08-14 cs.CV cs.AI 62%

Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy

Hao Yu, Rupayan Mallick, Margrit Betke, Sarah Adel Bargal

机构 * Boston University(波士顿大学) Georgetown University(乔治城大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09225 2025-08-14 eess.IV cs.AI cs.CV 62%

AMRG: Extend Vision Language Models for Automatic Mammography Report Generation

Nak-Jun Sung, Donghyun Lee, Bo Hwa Choi, Chae Jung Park

机构 * Research Institute, National Cancer Center Korea(韩国国家癌症中心研究所) Department of Radiology, National Cancer Center Korea(韩国国家癌症中心放射科)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09177 2025-08-14 eess.IV cs.AI cs.CV 62%

Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation

Xuanru Zhou, Cheng Li, Shuqiang Wang, Ye Li, Tao Tan, Hairong Zheng, Shanshan Wang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07146 2025-08-12 cs.CV cs.AI 62%

Intention-Aware Diffusion Model for Pedestrian Trajectory Prediction

Yu Liu, Zhijie Liu, Xiao Ren, You-Fu Li, He Kong

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏