arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-04 至 2025-11-04 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 14 篇

2510.24770 2025-11-04 eess.IV cs.AI cs.CV 84%

DMVFC: Deep Learning Based Functionally Consistent Tractography Fiber Clustering Using Multimodal Diffusion MRI and Functional MRI

Bocheng Guo, Jin Wang, Yijie Li, Junyi Wang, Mingyu Gao, Puming Feng, Yuqian Chen, Jarrett Rushmore, Nikos Makris, Yogesh Rathi, Lauren J O'Donnell, Fan Zhang

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01163 2025-11-04 cs.CV 83%

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation

Yongyuan Liang, Wei Chow, Feng Li, Ziqiao Ma, Xiyao Wang, Jiageng Mao, Jiuhai Chen, Jiatao Gu, Yue Wang, Furong Huang

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of Pennsylvania(宾夕法尼亚大学) The Hong Kong University of Science and Technology(香港科学与技术大学) University of Michigan(密歇根大学) University of Southern California(南加州大学)

专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Project Page: https://roverbench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19755 2025-11-04 cs.LG cs.AI cs.CV 81%

A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation

Jiacheng Liu, Xinyu Wang, Yuqi Lin, Zhikai Wang, Peiru Wang, Peiliang Cai, Qinming Zhou, Zhengan Yan, Zexuan Yan, Zhengyi Shi, Chang Zou, Yue Ma, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments 22 pages,2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01374 2025-11-04 cs.LG 78%

Learning Intractable Multimodal Policies with Reparameterization and Diversity Regularization

Ziqi Wang, Jiashun Liu, Ling Pan

机构 * Hong Kong University of Science and Technology(香港理工大学)

专题命中 多模态生成 :multimodal(title,abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24086 2025-11-04 cs.CV cs.CL 73%

MotionGPT3: Human Motion as a Second Modality

Bingfan Zhu, Biao Jiang, Sunyi Wang, Shixiang Tang, Tao Chen, Linjie Luo, Youyi Zheng, Xin Chen

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学) ByteDance(字节跳动) The Chinese University of HongKong(香港中文大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments 26 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12704 2025-11-04 cs.CV cs.MM 73%

SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding

Qianqian Sun, Jixiang Luo, Dell Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence(TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) The University of Hongkong(香港大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01593 2025-11-04 cs.CV 70%

Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation

Yizhu Chen, Chen Ju, Zhicheng Wang, Shuai Xiao, Xu Chen, Jinsong Lan, Xiaoyong Zhu, Ying Chen

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) Peking University(北京大学)

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14905 2025-11-04 eess.IV cs.AI cs.CL 62%

Structural Entities Extraction and Patient Indications Incorporation for Chest X-ray Report Generation

Kang Liu, Zhuoqi Ma, Xiaolu Kang, Zhusi Zhong, Zhicheng Jiao, Grayson Baird, Harrison Bai, Qiguang Miao

机构 * School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Xi'an Key Laboratory of Big Data and Intelligent Vision(西安大数据与智能视觉重点实验室) Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University(教育部协同智能系统重点实验室) Warren Alpert Medical School, Brown University(布朗大学沃伦·阿尔珀特医学院) School of Electronic Engineering, Xidian University(西安电子科技大学电子工程学院) Department of Radiology and Radiological Sciences, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院放射科)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments The code is available at https://github.com/mk-runner/SEI-Temp or https://github.com/mk-runner/SEI

Journal ref Medical Image Computing and Computer Assisted Intervention (MICCAI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00362 2025-11-04 cs.CV cs.AI cs.GR 62%

Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery

Momen Khandoker Ope, Akif Islam, Mohd Ruhul Ameen, Abu Saleh Musa Miah, Md Rashedul Islam, Jungpil Shin

机构 * University of Rajshahi(拉贾沙希大学) Marshall University(马歇尔大学) University of Aizu(御所大学) University of Asia Pacific(亚洲太平洋大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 6 Pages, 4 figures, 2 Tables, Submitted to ICECTE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00107 2025-11-04 cs.CV cs.AI cs.IR 62%

AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency

Piyushkumar Patel

机构 * Microsoft(微软)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01645 2025-11-04 cs.CV 57%

Enhancing Diffusion-based Restoration Models via Difficulty-Adaptive Reinforcement Learning with IQA Reward

Xiaogang Xu, Ruihang Chu, Jian Wang, Kun Zhou, Wenjie Shu, Harry Yang, Ser-Nam Lim, Hao Chen, Liang Lin

机构 * The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Snap Research Shenzhen University(深圳大学) HKUST(香港科技大学) University of Central Florida(佛罗里达大学) UC Davis(加州大学戴维斯分校) Sun Yat-Sen University(孙中山大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16800 2025-11-04 cs.CV 57%

Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception

Jiajing Lin, Zhenzhong Wang, Dejun Xu, Shu Jiang, YunPeng Gong, Min Jiang

机构 * School of Informatics, Xiamen University(厦门大学信息学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025. Project Page: https://jiajinglin.github.io/Phys4DGen

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00344 2025-11-04 cs.CV 57%

Federated Dialogue-Semantic Diffusion for Emotion Recognition under Incomplete Modalities

Xihang Qiu, Jiarong Cheng, Yuhao Fang, Wanpeng Zhang, Yao Lu, Ye Zhang, Chun Li

机构 * Shenzhen MSU-BIT University(深圳MSU-BIT大学) Beijing Institude of Technology(北京理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15069 2025-11-04 stat.CO stat.ME stat.ML 50%

Sampling by averaging: A multiscale approach to score estimation

Paula Cordero-Encinar, Andrew B. Duncan, Sebastian Reich, O. Deniz Akyildiz

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏