arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-26 至 2025-08-26 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 14 篇

2508.17376 2025-08-26 cs.LG cs.CV 83%

ShaLa: Multimodal Shared Latent Space Modelling

Jiali Cui, Yan-Ying Chen, Yanxia Zhang, Matthew Klenk

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16930 2025-08-26 eess.AS cs.CV cs.SD 81%

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Sizhe Shan, Qiulin Li, Yutao Cui, Miles Yang, Yuehai Wang, Qun Yang, Jin Zhou, Zhao Zhong

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03001 2025-08-26 cs.CV cs.MM 81%

One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning

Hao Sun, Yu Song, Jiaqing Liu, Jihong Hu, Yen-Wei Chen, Lanfen Lin

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Information Science and Engineering, Ritsumeikan University(立命馆大学信息科学与工程学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17614 2025-08-26 cs.CV 79%

JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on

Aowen Wang, Wei Li, Hao Luo, Mengxing Ao, Chenyu Zhu, Xinyang Li, Fan Wang

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(汇安实验室) Zhejiang University(浙江大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17199 2025-08-26 cs.CV 79%

MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling

Hyeyeon Kim, Sungwoo Han, Jingun Kwon, Hidetaka Kamigaito, Manabu Okumura

机构 * Chungnam National University(Chungnam 国立大学) Nara Institute of Science and Technology (NAIST)(Nara 科学技术研究所) Institute of Science Tokyo(东京科学研究所)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16763 2025-08-26 cs.CV 79%

WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation

Rabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow Mila Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) École de Technologie Supérieure (ETS)(高等技术学院) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments This paper has been accepted to the EMNLP 2025 main conference. Check the project page here: https://webmmu-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17718 2025-08-26 cs.CV cs.AI 73%

Instant Preference Alignment for Text-to-Image Diffusion Models

Yang Li, Songlin Yang, Xiaoxuan Han, Wei Wang, Jing Dong, Yueming Lyu, Ziyu Xue

机构 * New Laboratory of Pattern Recognition, CASIA(模式识别新实验室,中国科学院自动化研究所) The Hong Kong University of Science and Technology(香港科技大学) Nanjing university(南京大学) Academy of Broadcasting Science, NRTA(广播科学研究院,国家广播电视总局)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17342 2025-08-26 cs.GR cs.CV cs.MM cs.SD 62%

DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions

Hengyuan Zhang, Zhe Li, Xingqun Qi, Mengze Li, Muyi Sun, Man Zhang, Sirui Han

机构 * Peking University(北京大学) The Hong Kong University of Science and Technology(香港科技大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.MM

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17062 2025-08-26 cs.CV cs.AI 62%

SSG-Dit: A Spatial Signal Guided Framework for Controllable Video Generation

Peng Hu, Yu Gu, Liang Luo, Fuji Ren

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(计算机科学与工程学院,电子科技大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18213 2025-08-26 cs.CV 57%

Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance

Ayce Idil Aytekin, Helge Rhodin, Rishabh Dabral, Christian Theobalt

机构 * Max Planck Institute for Informatics and Saarland University(马克斯·普朗克研究所(信息学)和萨尔兰大学) Bielefeld University(比勒菲尔德大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Project page: https://aidilayce.github.io/FollowMyHold-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08402 2025-08-26 cs.CV 57%

V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object Detection

Xun Huang, Jinlong Wang, Qiming Xia, Siheng Chen, Bisheng Yang, Xin Li, Cheng Wang, Chenglu Wen

机构 * Fujian Key Laboratory of Sensing and Computing for Smart Cities, Xiamen University, China(福建智能城市感知与计算重点实验室,厦门大学) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China(多媒体可信感知与高效计算重点实验室,中国教育部,厦门大学) Zhongguancun Academy(中关村学院) Shanghai Jiao Tong University(上海交通大学) Wuhan University(武汉大学) Texas A&M University(德克萨斯A&M大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16643 2025-08-26 cs.LG cs.AI 57%

From Classical Probabilistic Latent Variable Models to Modern Generative AI: A Unified Perspective

Tianhua Chen

机构 * School of Computing and Engineering University of Huddersfield(计算与工程学院赫德斯菲尔德大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments This is a substantially improved and expanded version of an earlier manuscript hosted on SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5244929

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12835 2025-08-26 cs.CV 57%

DiffS-NOCS: 3D Point Cloud Reconstruction through Coloring Sketches to NOCS Maps Using Diffusion Models

Di Kong, Qianhui Wan

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Zhongguancun Academy(中关村学院) Beijing Normal University(北京师范大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00210 2025-08-26 cs.LG cs.CE cs.SY eess.SY 50%

Generative Machine Learning in Adaptive Control of Dynamic Manufacturing Processes: A Review

Suk Ki Lee, Hyunwoong Ko

机构 * School of Manufacturing Systems and Networks, Arizona State University(制造系统与网络学院,亚利桑那州立大学)

专题命中 多模态生成 :multimodal(abstract)

Comments 12 pages, 1 figure, 1 table. This paper has been accepted for publication in the proceedings of ASME IDETC-CIE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏