arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-18 至 2025-11-18 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 15 篇

2511.13647 2025-11-18 cs.CV 88%

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

Chunshi Wang, Junliang Ye, Yunhan Yang, Yang Li, Zizhuo Lin, Jun Zhu, Zhuo Chen, Yawei Luo, Chunchao Guo

机构 * Zhejiang University(浙江大学) Tencent Hunyuan(腾讯文言) Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13243 2025-11-18 cs.LG cs.AI cs.CV 84%

Uncovering and Mitigating Transient Blindness in Multimodal Model Editing

Xiaoqi Han, Ru Li, Ran Yi, Hongye Tan, Zhuomin Liang, Víctor Gutiérrez-Basulto, Jeff Z. Pan

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at AAAI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13078 2025-11-18 cs.LG eess.AS eess.IV 79%

A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning

Liuyi Jin, Pasan Gunawardena, Amran Haroon, Runzhi Wang, Sangwoo Lee, Radu Stoleru, Michael Middleton, Zepeng Huo, Jeeeun Kim, Jason Moats

机构 * Computer Engineering, Texas A\&M University 3 Texas A\&M University Emergency Medical Services (EMS), 4 Biomedical Data Science, Stanford University, 5 Texas A\&M School of Public Health

专题命中 多模态生成 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23639 2025-11-18 cs.LG cs.AI q-bio.QM 79%

Integrating Genomics into Multimodal EHR Foundation Models

Jonathan Amar, Edward Liu, Alessandra Breschi, Liangliang Zhang, Pouya Kheradpour, Sylvia Li, Lisa Soleymani Lehmann, Alessandro Giulianelli, Matt Edwards, Yugang Jia, David Nola, Raghav Mani, Pankaj Vats, Jesse Tetreault, T. J. Chen, Cory Y. McLean

机构 * Verily Life Sciences(Verily生命科学公司) Nvidia(英伟达公司) Google(谷歌公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12280 2025-11-18 cs.CV cs.CL cs.LG 74%

D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs

Shuochen Chang, Xiaofeng Zhang, Qingyang Liu, Li Niu

机构 * Project leader(项目负责人)

专题命中 多模态生成 :MLLM(abstract,comments);multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by AAAI Conference on Artificial Intelligence (AAAI) 2026. Code available at https://github.com/bcmi/D3ToM-Diffusion-MLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12072 2025-11-18 cs.MM cs.AI cs.SD 73%

ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation

Jiahui Sun, Weining Wang, Mingzhen Sun, Yirong Yang, Xinxin Zhu, Jing Liu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beihang University(北航大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12363 2025-11-18 cs.CV 70%

Explainable AI-Generated Image Detection RewardBench

Michael Yang, Shijian Deng, William T. Doan, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) University of Notre Dame(诺特大学) Stony Brook University(石溪大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12851 2025-11-18 cs.CL cs.AI 62%

NeuroLex: A Lightweight Domain Language Model for EEG Report Understanding and Generation

Kang Yin, Hye-Bin Shin

机构 * Dept. of Artificial Intelligence Korea University Seoul, Republic of Korea(人工智能系韩国大学首尔共和国韩国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03457 2025-11-18 cs.GR cs.CV cs.SD eess.AS 62%

READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation

Haotian Wang, Yuzhe Weng, Jun Du, Haoran Xu, Xiaoyan Wu, Shan He, Bing Yin, Cong Liu, Jianqing Gao, Qingfeng Liu

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments Project page: https://readportrait.github.io/READ/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15217 2025-11-18 cs.SD cs.AI cs.LG cs.MM 62%

DRAGON: Distributional Rewards Optimize Diffusion Generative Models

Yatong Bai, Jonah Casebeer, Somayeh Sojoudi, Nicholas J. Bryan

机构 * University of California, Berkeley(加州大学伯克利分校) Adobe Research(Adobe研究)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI、cs.MM

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00998 2025-11-18 cs.CV cs.AI 62%

FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image Translation

Xiang Gao, Jiaying Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted conference paper of ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13309 2025-11-18 cs.CV 57%

DriveLiDAR4D: Sequential and Controllable LiDAR Scene Generation for Autonomous Driving

Kaiwen Cai, Xinze Liu, Xia Zhou, Hengtong Hu, Jie Xiang, Luyao Zhang, Xueyang Zhang, Kun Zhan, Yifei Zhan, Xianpeng Lang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12428 2025-11-18 cs.CV 57%

RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning

Jingqi Xu, Jingxi Lu, Chenghao Li, Sreetama Sarkar, Souvik Kundu, Peter A. Beerel

机构 * University of Southern California(南加州大学) Intel Labs(英特尔实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12040 2025-11-18 cs.CV 57%

SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images

Xinyuan Hu, Changyue Shi, Chuxiao Yang, Minghao Chen, Jiajun Ding, Tao Wei, Chen Wei, Zhou Yu, Min Tan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments AAAI2026-Oral. Project Page: https://xinyuanhu66.github.io/SRSplat/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13186 2025-11-18 cs.LG cs.SY eess.SY 50%

DiffFP: Learning Behaviors from Scratch via Diffusion-based Fictitious Play

Akash Karthikeyan, Yash Vardhan Pant

机构 * Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系)

专题命中 多模态生成 :multimodal(abstract)

Comments Initial results presented at the IJCAI 2025 Workshop on User-Aligned Assessment of Adaptive AI Systems. Project page: https://aku02.github.io/projects/difffp/

详情

展开后加载摘要…

URL PDF HTML 收藏