arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-05 至 2025-08-05 共收录 17 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 17 篇

2508.01615 2025-08-05 cs.LG cs.AI 83%

TCDiff: Triplex Cascaded Diffusion for High-fidelity Multimodal EHRs Generation with Incomplete Clinical Data

Yandong Yan, Chenxi Li, Yu Huang, Dexuan Xu, Jiaqi Zhu, Zhongyan Chai, Huamin Zhang

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) National Engineering Research Center for Software Engineering, Peking University(软件工程国家工程研究中心) Institute of Software, Chinese Academy of Science(中国科学院软件研究所) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) Institute of Basic Theory of Chinese Medicine, China Academy of Chinese Medical Sciences(中国中医科学院中医基础理论研究所)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16326 2025-08-05 cs.LG 82%

ChemMLLM: Chemical Multimodal Large Language Model

Qian Tan, Dongzhan Zhou, Peng Xia, Wanhao Liu, Wanli Ouyang, Lei Bai, Yuqiang Li, Tianfan Fu

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01058 2025-08-05 eess.IV cs.CV 79%

ReCoSeg++:Extended Residual-Guided Cross-Modal Diffusion for Brain Tumor Segmentation

Sara Yavari, Rahul Nitin Pandya, Jacob Furst

机构 * School of Computing, DePaul University(计算学院,德保罗大学)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01577 2025-08-05 eess.IV cs.CV 79%

Tractography-Guided Dual-Label Collaborative Learning for Multi-Modal Cranial Nerves Parcellation

Lei Xie, Junxiong Huang, Yuanjing Feng, Qingrun Zeng

机构 * Zhejiang University of Technology(浙江工业大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21562 2025-08-05 cs.CL cs.AI cs.AR 76%

FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction

Jun Yin, Pengyu Zeng, Jing Zhong, Peilin Li, Miao Zhang, Ran Luo, Shuai Lu

专题命中 多模态生成 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02362 2025-08-05 cs.CV cs.AI 62%

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

Xu Wang, Shengeng Tang, Fei Wang, Lechao Cheng, Dan Guo, Feng Xue, Richang Hong

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00443 2025-08-05 cs.CV 57%

SDMatte: Grafting Diffusion Models for Interactive Matting

Longfei Huang, Yu Liang, Hao Zhang, Jinwei Chen, Wei Dong, Lunde Chen, Wanyu Liu, Bo Li, Peng-Tao Jiang

机构 * Shanghai University(上海大学) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted at ICCV 2025, 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17349 2025-08-05 cs.CV cs.IR 57%

DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition

Yiyan Xu, Wuqiang Zheng, Wenjie Wang, Fengbin Zhu, Xinting Hu, Yang Zhang, Fuli Feng, Tat-Seng Chua

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication in ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01778 2025-08-05 cs.CV cs.RO 57%

DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online HD Map Diffusion

Zhigang Sun, Yiru Wang, Anqing Jiang, Shuo Wang, Yu Gao, Yuwen Heng, Shouyi Zhang, An He, Hao Jiang, Jinhao Chai, Zichong Gu, Wang Jijun, Shichen Tang, Lavdim Halilaj, Juergen Luettin, Hao Sun

机构 * Bosch Corporate Research, Bosch (China) Investment Ltd.(博世企业研究院、博世(中国)投资有限公司) School of Communication and Information Engineering, Shanghai University(上海大学通信与信息工程学院) Shanghai Jiaotong University(上海交通大学) AIR, Tsinghua University(清华大学人工智能研究院) Robert Bosch GmbH(罗伯特·博世集团)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01761 2025-08-05 cs.LG cs.AI 57%

Semantically-Guided Inference for Conditional Diffusion Models: Enhancing Covariate Consistency in Time Series Forecasting

Rui Ding, Hanyang Meng, Zeyang Zhang, Jielong Yang

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02485 2025-08-05 cs.AI cs.CE 57%

Generative AI as a Pillar for Predicting 2D and 3D Wildfire Spread: Beyond Physics-Based Models and Traditional Deep Learning

Haowen Xu, Sisi Zlatanova, Ruiyu Liang, Ismet Canbulat

机构 * GRID, School of Built Environment, UNSW Sydney, NSW 2052 Australia(GRID,环境建筑学院,新南威尔士大学悉尼分校) School of Minerals and Energy Resources Engineering, UNSW Sydney, NSW 2052 Australia(矿物与能源资源工程学院,新南威尔士大学悉尼分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11935 2025-08-05 cs.CL 57%

ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing

Xuanle Zhao, Xuexin Liu, Haoyue Yang, Xianzhen Luo, Fanhu Zeng, Jianling Li, Qi Shi, Chi Chen

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Harbin Institute of Technology(哈尔滨工业大学) Tianjin University(天津大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments Accepted by ACL2025 Findings, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13490 2025-08-05 cs.CV 57%

Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing

Joowon Kim, Ziseok Lee, Donghyeon Cho, Sanghyun Jo, Yeonsung Jung, Kyungsu Kim, Eunho Yang

机构 * KAIST(韩国科学技术院) Seoul National University(首尔国立大学) OGQ

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07580 2025-08-05 cs.CV 57%

InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior

Chenguo Lin, Yuchen Lin, Panwang Pan, Xuanyang Zhang, Yadong Mu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) PICO AI group, ByteDance(字节跳动PICO AI团队)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to T-PAMI 2025. This paper is an extension of ICLR 2024 "InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior". arXiv admin note: substantial text overlap with arXiv:2402.04717

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01130 2025-08-05 cs.CV 57%

Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models

Bicheng Xu, Qi Yan, Renjie Liao, Lele Wang, Leonid Sigal

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02117 2025-08-05 eess.SP 50%

Scoring ISAC: Benchmarking Integrated Sensing and Communications via Score-Based Generative Modeling

Lin Chen, Chang Cai, Huiyuan Yang, Xiaojun Yuan, Ying-Jun Angela Zhang

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01975 2025-08-05 cs.LG stat.ML 50%

Diffusion models for inverse problems

Hyungjin Chung, Jeongsol Kim, Jong Chul Ye

机构 * EverEx KAIST(韩国科学技术院)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏