arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2509.11406 2025-09-16 cs.CV 57%

No Modality Left Behind: Dynamic Model Generation for Incomplete Medical Data

Christoph Fürböck, Paul Weiser, Branko Mitic, Philipp Seeböck, Thomas Helbich, Georg Langs

机构 * Computational Imaging Research Lab(计算成像研究实验室) Department for Biomedical Imaging and Image-guided Therapy(生物医学成像与影像引导治疗部门) Medical University of Vienna(维也纳医学大学) Comprehensive Center for Artificial Intelligence in Medicine(医学人工智能综合中心) Christian Doppler Laboratory for Machine Learning Driven Precision Imaging(机器学习驱动精准成像的克里斯蒂安·多普勒实验室) Department of Biomedical Imaging and Image-guided Therapy(生物医学成像与影像引导治疗部门) Athinoula A. Martinos Center for Biomedical Imaging(阿提诺拉·A·马丁努斯生物医学成像中心) Massachusetts General Hospital(麻省总医院) Harvard Medical School(哈佛医学院) Department of Radiology(放射科) Division of General and Pediatric Radiology(普通和儿童放射科部门)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted at MICCAI2025 ML-CDS Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10873 2025-09-16 cs.MM 57%

Automated Radiology Report Generation Based on Topic-Keyword Semantic Guidance

Jing Xiao, Hongfei Liu, Ruiqi Dong, Jimin Liu, Haoyong Yu

专题命中 多模态生成 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09717 2025-09-15 cs.SD cs.LG eess.AS 57%

Testing chatbots on the creation of encoders for audio conditioned image generation

Jorge E. León, Miguel Carrasco

专题命中 多模态生成 :image-text(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06932 2025-09-11 cs.RO cs.CV 57%

LLaDA-VLA: Vision Language Diffusion Action Models

Yuqing Wen, Hebei Li, Kefan Gu, Yucheng Zhao, Tiancai Wang, Xiaoyan Sun

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学) Dexmal Project Page(Dexmal项目页)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04242 2025-09-11 cs.LG cs.CV 57%

Task-based Loss Functions in Computer Vision: A Comprehensive Review

Omar Elharrouss, Yasir Mahmood, Yassine Bechqito, Mohamed Adel Serhani, Elarbi Badidi, Jamal Riffi, Hamid Tairi

机构 * Department of Computer Science and Software Engineering, College of Information Technology, United Arab Emirates University.(计算机科学与软件工程系,信息科技学院,阿联酋大学) Department of Information Systems, College of Computing and Informatics, University of Sharjah, Sharjah, United Arab Emirates(信息系统系,计算与信息学院,沙迦大学) Department of Informatics, Faculty of Sciences Dhar El Mahraz, Sidi Mohamed Ben Abdellah University, Fez, Morocco(信息学系,达尔·埃尔·马哈勒兹学院,西迪·莫哈梅德·本·阿卜杜勒拉赫曼大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04444 2025-09-10 cs.CV 57%

One Flight Over the Gap: A Survey from Perspective to Panoramic Vision

Xin Lin, Xian Ge, Dizhe Zhang, Zhaoliang Wan, Xianshun Wang, Xiangtai Li, Wenjie Jiang, Bo Du, Dacheng Tao, Ming-Hsuan Yang, Lu Qi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://insta360-research-team.github.io/Survey-of-Panorama/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03726 2025-09-09 cs.LG cs.AI q-bio.BM 57%

Diffusion on language model encodings for protein sequence generation

Viacheslav Meshchaninov, Pavel Strashnov, Andrey Shevtsov, Fedor Nikolaev, Nikita Ivanisenko, Olga Kardymon, Dmitry Vetrov

机构 * Constructor University, Bremen, Germany(Constructor大学,不来梅,德国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04269 2025-09-05 cs.CV 57%

TauGenNet: Plasma-Driven Tau PET Image Synthesis via Text-Guided 3D Diffusion Models

Yuxin Gong, Se-in Jang, Wei Shao, Yi Su, Kuang Gong

机构 * J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida(J. Crayton Pruitt家族生物医学工程系,佛罗里达大学) Department of Radiology & Biomedical Imaging, Yale University(放射学与生物医学成像系,耶鲁大学) Department of Medicine, University of Florida(医学系,佛罗里达大学) Banner Alzheimer’s Institute(Banner阿尔茨海默症研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 4 figures, submitted to IEEE Transactions on Radiation and Plasma Medical Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03550 2025-09-05 cs.AI 57%

Diffusion-RL Based Air Traffic Conflict Detection and Resolution Method

Tonghe Li, Jixin Liu, Weili Zeng, Hao Jiang

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 59 pages,13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16507 2025-09-05 cs.CV cs.LG 57%

Straighter Flow Matching via a Diffusion-Based Coupling Prior

Siyu Xing, Jie Cao, Huaibo Huang, Haichao Shi, Xiao-Yu Zhang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14906 2025-09-04 eess.IV cs.CV 57%

FetalFlex: Anatomy-Guided Diffusion Model for Flexible Control on Fetal Ultrasound Image Synthesis

Yaofei Duan, Tao Tan, Zhiyuan Zhu, Yuhao Huang, Yuanji Zhang, Rui Gao, Patrick Cheong-Iao Pang, Xinru Gao, Guowei Tao, Xiang Cong, Zhou Li, Lianying Liang, Guangzhi He, Linliang Yin, Xuedong Deng, Xin Yang, Dong Ni

机构 * organization= Faculty of Applied Sciences, Macao Polytechnic University , city= Macao , country= China organization= National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University , city= Shenzhen , state= Guangdong , country= China organization= Shenzhen RayShape Medical Technology Co., Ltd , city= Shenzhen , state= Guangdong , country= China organization= Department of Ultrasound, Shenzhen Guangming District People’s Hospital , city= Shenzhen , state= Guangdong , country= China organization= Jinan University , city= Guangzhou , state= Guangdong , country= China organization= Center for Medical Ultrasound, The Affiliated Suzhou Hospital of Nanjing Medical University, Suzhou Municipal Hospital, Gusu School, Nanjing Medical University , city= Suzhou , state= Jiangsu , country= China organization= Medical Ultrasound Image Computing (MUSIC) Laboratory, Shenzhen University , city= Shenzhen , state= Guangdong , country= China organization= Qilu Hospital of Shandong University , city= Jinan , state= Shandong , country= China organization= Northwest Women \& Children Hospital , city= Xian , state= Shaanxi , country= China

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 18 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12470 2025-09-03 cs.CV 57%

SC-Diff: 3D Shape Completion with Latent Diffusion Models

Simon Schaefer, Juan D. Galvis, Xingxing Zuo, Stefan Leutengger

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) ETH Zurich(苏黎世联邦理工学院) MBZUAI(穆桑比克大学人工智能研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00428 2025-09-03 cs.CV 57%

Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

Xuechao Zou, Shun Zhang, Xing Fu, Yue Li, Kai Li, Yushe Cao, Congyan Lang, Pin Tao, Junliang Xing

机构 * Beijing Jiaotong University(北京交通大学) Ant Group(蚂蚁集团) Qinghai University(青海大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18512 2025-09-03 physics.optics cs.CL 57%

Designing across domains with declarative thinking: Insights from the 96-Eyes ptychographic imager project

Antony C Chan

机构 * Consultant, high-throughput microscopy and hardware-accelerated algorithms(咨询顾问,高通量显微镜和硬件加速算法)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments Minor changes: resolve HTML rendering issues of sideways tables; Code listing in dark mode. Cite three more journal articles

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05304 2025-09-03 cs.LG cs.CV 57%

Gaussian Mixture Flow Matching Models

Hansheng Chen, Kai Zhang, Hao Tan, Zexiang Xu, Fujun Luan, Leonidas Guibas, Gordon Wetzstein, Sai Bi

机构 * Stanford University(斯坦福大学) Adobe Research(Adobe研究)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICML 2025. Code: https://github.com/Lakonik/GMFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19852 2025-08-29 cs.CV 57%

Ego-centric Predictive Model Conditioned on Hand Trajectories

Binjie Zhang, Mike Zheng Shou

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Code: github.com/showlab/Ego-PM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20020 2025-08-28 cs.CV 57%

GS: Generative Segmentation via Label Diffusion

Yuhao Chen, Shubin Chen, Liang Lin, Guangrun Wang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19508 2025-08-28 cs.RO cs.CV 57%

DATR: Diffusion-based 3D Apple Tree Reconstruction Framework with Sparse-View

Tian Qiu, Alan Zoubi, Yiyuan Lin, Ruiming Du, Lailiang Cheng, Yu Jiang

机构 * School of Electrical and Computer Engineering, Cornell University(电气与计算机工程系,康奈尔大学) Sibley School of Mechanical and Aerospace Engineering, Cornell University(机械与航空航天工程系,康奈尔大学) School of Biological and Environmental Engineering, Cornell University(生物与环境工程系,康奈尔大学) School of Integrative Plant Science, Cornell University(整合植物科学系,康奈尔大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15842 2025-08-28 cs.CV cs.GR 57%

DiffArtist: Towards Structure and Appearance Controllable Image Stylization

Ruixiang Jiang, Changwen Chen

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACM MM 2025, Homepage: https://DiffusionArtist.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14874 2025-08-28 cs.CV 57%

TraceNet: Segment one thing efficiently

Mingyuan Wu, Zichuan Liu, Haozhen Zheng, Hongpeng Guo, Bo Chen, Xin Lu, Klara Nahrstedt

机构 * Coordinated Science Laboratory, University of Illinois at Urbana-Champaign, Champaign, USA(伊利诺伊大学厄巴纳-香槟分校协调科学实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Best Student Paper in IEEE MIPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15761 2025-08-27 cs.CV 57%

Waver: Wave Your Way to Lifelike Video Generation

Yifu Zhang, Hao Yang, Yuqi Zhang, Yifei Hu, Fengda Zhu, Chuang Lin, Xiaofeng Mei, Yi Jiang, Bingyue Peng, Zehuan Yuan

机构 * Bytedance Waver Team(字节跳动Waver团队)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06905 2025-08-27 cs.CV 57%

MultiRef: Controllable Image Generation with Multiple Visual References

Ruoxi Chen, Dongping Chen, Siyuan Wu, Sinan Wang, Shiyun Lang, Petr Sushko, Gaoyang Jiang, Yao Wan, Ranjay Krishna

机构 * Zhejiang Wanli University(浙江万里大学) University of Washington(华盛顿大学) Huazhong University of Science and Technology(华中科技大学) Allen Institute for AI(人工智能研究院)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted to ACM MM 2025 Datasets

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18213 2025-08-26 cs.CV 57%

Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance

Ayce Idil Aytekin, Helge Rhodin, Rishabh Dabral, Christian Theobalt

机构 * Max Planck Institute for Informatics and Saarland University(马克斯·普朗克研究所(信息学)和萨尔兰大学) Bielefeld University(比勒菲尔德大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Project page: https://aidilayce.github.io/FollowMyHold-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08402 2025-08-26 cs.CV 57%

V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object Detection

Xun Huang, Jinlong Wang, Qiming Xia, Siheng Chen, Bisheng Yang, Xin Li, Cheng Wang, Chenglu Wen

机构 * Fujian Key Laboratory of Sensing and Computing for Smart Cities, Xiamen University, China(福建智能城市感知与计算重点实验室,厦门大学) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China(多媒体可信感知与高效计算重点实验室,中国教育部,厦门大学) Zhongguancun Academy(中关村学院) Shanghai Jiao Tong University(上海交通大学) Wuhan University(武汉大学) Texas A&M University(德克萨斯A&M大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16643 2025-08-26 cs.LG cs.AI 57%

From Classical Probabilistic Latent Variable Models to Modern Generative AI: A Unified Perspective

Tianhua Chen

机构 * School of Computing and Engineering University of Huddersfield(计算与工程学院赫德斯菲尔德大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments This is a substantially improved and expanded version of an earlier manuscript hosted on SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5244929

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12835 2025-08-26 cs.CV 57%

DiffS-NOCS: 3D Point Cloud Reconstruction through Coloring Sketches to NOCS Maps Using Diffusion Models

Di Kong, Qianhui Wan

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Zhongguancun Academy(中关村学院) Beijing Normal University(北京师范大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14879 2025-08-25 cs.GR cs.CV 57%

MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds

Bingquan Dai, Li Ray Luo, Qihong Tang, Jie Wang, Xinyu Lian, Hao Xu, Minghan Qin, Xudong Xu, Bo Dai, Haoqian Wang, Zhaoyang Lyu, Jiangmiao Pang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Harbin Institute of Technology(哈尔滨工业大学) Beijing Institute of Technology(北京理工大学) AI Thrust, HKUST(GZ)(AI thrust, 香港科技大学(广州))

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22306 2025-08-22 cs.LG cs.AI 57%

Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer

Zehua Chen, Yuyang Miao, Liyuan Wang, Luyun Fan, Danilo P. Mandic, Jun Zhu

机构 * Department of Computer Science and Technology, Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint Center for ML, Tsinghua University, Beijing, China(计算机科学与技术系、人工智能研究院、BNRist中心、THBI实验室、清华大学-博世联合机器学习中心、清华大学、北京,中国) Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing, China(心理学与认知科学系、清华大学、北京,中国) Department of Electrical and Electronic Engineering, Imperial College London, London, United Kingdom(电子与电气工程系、伦敦帝国理工学院、伦敦,英国) Beijing Anzhen Hospital of Capital Medical University, Beijing Institute of Heart Lung and Blood Vessel Diseases, Chinese Institutes for Medical Research, Beijing, China(首都医科大学北京安贞医院、北京心肺血管疾病研究院、中国医学研究院、北京,中国)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14718 2025-08-21 cs.CL 57%

The Digital Sous Chef -- A Comparative Study on Fine-Tuning Language Models for Recipe Generation

Shubham Pundhir, Ganesh Bagler

机构 * Indraprastha Institute of Information Technology(印度理工学院信息技术研究所)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments 8 pages, 4 figures. Code is available at: https://github.com/shubh-iiit/RecipeGPT2-Your-Own-AI-Chef

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14405 2025-08-21 cs.CV 57%

CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities

Yue Gong, Shanyuan Liu, Liuzhuozheng Li, Jian Zhu, Bo Cheng, Liebucha Wu, Xiaoyu Wu, Yuhang Ma, Dawei Leng, Yuhui Yin

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏