arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2508.14393 2025-08-21 cs.CV 57%

Img2ST-Net: Efficient High-Resolution Spatial Omics Prediction from Whole Slide Histology Images via Fully Convolutional Image-to-Image Learning

Junchao Zhu, Ruining Deng, Junlin Guo, Tianyuan Yao, Juming Xiong, Chongyu Qu, Mengmeng Yin, Yu Wang, Shilin Zhao, Haichun Yang, Daguang Xu, Yucheng Tang, Yuankai Huo

机构 * Department of Computer Science, Vanderbilt University(计算机科学系,范德比尔特大学) Weill Cornell Medicine(韦尔·科恩医学中心) Department of Electrical and Computer Engineering, Vanderbilt University(电气与计算机工程系,范德比尔特大学) Department of Biostatistics, Vanderbilt University Medical Center(生物统计学系,范德比尔特大学医学中心) NVIDIA

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14359 2025-08-21 cs.CV 57%

Taming Transformer for Emotion-Controllable Talking Face Generation

Ziqi Zhang, Cheng Deng

机构 * Xidian University(西安电子科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12512 2025-08-19 cs.CV 57%

LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models

Krishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICIP 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12396 2025-08-19 cs.CV 57%

DeCoT: Decomposing Complex Instructions for Enhanced Text-to-Image Generation with Large Language Models

Xiaochuan Lin, Xiangyong Chen, Xuan Li, Yichen Su

机构 * Henan Polytechnic University(河南理工大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12361 2025-08-19 cs.LG cs.AI math.ST stat.TH 57%

Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models

Xun Su, Jianming Huang, Yang Yusen, Zhongxi Fang, Hiroyuki Kasai

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09693 2025-08-19 cs.CV 57%

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments

Jiali Chen, Yujie Jia, Zihan Wu, Jinyu Yang, Jianpeng Chen, Xusen Hei, Jiayuan Xie, Yi Cai, Qing Li

机构 * South China University of Technology(南方科技大学) The Hong Kong Polytechnic University(香港理工大学) Key Laboratory of Big Data and Intelligent Robot Ministry of Education(教育部大数据与智能机器人重点实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11153 2025-08-18 cs.CV 57%

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction

Maoquan Zhang, Bisser Raytchev, Xiujuan Sun

机构 * Graduate School of Advanced Science and Engineering, Hiroshima University(Hiroshima大学研究生院) Department of Computer Science, Weifang University of Science and Technology(潍坊科技大学计算机科学系)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments The International Conference on Neural Information Processing (ICONIP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05908 2025-08-18 cs.CV 57%

GBR: Generative Bundle Refinement for High-fidelity Gaussian Splatting with Enhanced Mesh Reconstruction

Jianing Zhang, Yuchao Zheng, Ziwei Li, Qionghai Dai, Xiaoyun Yuan

机构 * College of future information technology, Fudan University(未来信息科技学院,复旦大学) School of Biomedical Engineering, Tsinghua University(生物医学工程学院,清华大学) Key Laboratory for Information Science of Electromagnetic Waves (MoE), Fudan University(电磁波信息科学重点实验室(MoE),复旦大学) Department of Automation, Tsinghua University(自动化系,清华大学) MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能重点实验室,人工智能研究院,上海交通大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10858 2025-08-15 cs.CV 57%

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

Harold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang, Ser-Nam Lim

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Project Page: https://haroldchen19.github.io/PhysHPO-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10280 2025-08-15 cs.CV 57%

High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance

Danyi Gao

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10013 2025-08-15 cs.CL 57%

Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis

Linqing Chen, Hanmeng Zhong, Wentao Wu, Weilei Wang

机构 * Linqing Chen, Hanmeng Zhong, Wentao Wu, Weilei Wang(作者)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01272 2025-08-15 cs.CV 57%

PromptSafe: Gated Prompt Tuning for Safe Text-to-Image Generation

Zonglei Jing, Xiao Yang, Xiaoqian Li, Siyuan Liang, Aishan Liu, Mingchuan Zhang, Xianglong Liu

机构 * Beihang University(北航) Beijing University of Posts and Telecommunications(北京邮电大学) Taishan University(泰山大学) Nanyang Technological University(南洋理工大学) Henan University of Science and Technology(河南科技大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01657 2025-08-14 cs.IR cs.CV 57%

RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation

Run Ling, Wenji Wang, Yuting Liu, Guibing Guo, Haowei Liu, Jian Lu, Quanwei Zhang, Yexing Xu, Shuo Lu, Yun Wang, Yihua Shao, Zhanjie Zhang, Ao Ma, Linying Jiang, Xingwei Wang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09444 2025-08-14 cs.RO cs.CV 57%

DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation

Haoxiang Shi, Xiang Deng, Zaijing Li, Gongwei Chen, Yaowei Wang, Liqiang Nie

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08891 2025-08-13 cs.CV 57%

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos

Chaoyi Wang, Yifan Yang, Jun Pei, Lijie Xia, Jianpo Liu, Xiaobing Yuan, Xinhan Di

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments This paper has been accepted by ICCV 2025 Workshop MMFM4

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08220 2025-08-12 cs.CV 57%

Learning User Preferences for Image Generation Model

Wenyi Mo, Ying Ba, Tianyu Zhang, Yalong Bai, Biye Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07540 2025-08-12 cs.CV 57%

CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts

Junuk Cha, Jihyeon Kim

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICCVW'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07225 2025-08-12 eess.IV cs.CV q-bio.QM 57%

HaDM-ST: Histology-Assisted Differential Modeling for Spatial Transcriptomics Generation

Xuepeng Liu, Zheng Jiang, Pinan Zhu, Hanyu Liu, Chao Li

机构 * University of Cambridge, UK(剑桥大学,英国) Northeastern University, Shenyang, China(东北大学,中国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 5 figures, includes comparisons with TESLA, HiStoGene, and iStar; submitted to arXiv 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16726 2025-08-12 cs.CV cs.LG 57%

EDiT: Efficient Diffusion Transformers with Linear Compressed Attention

Philipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick, Luca Morreale, Mehdi Noroozi, Alberto Gil Ramos, Sourav Bhattacharya

机构 * Samsung, AI Center Cambridge(三星人工智能中心剑桥)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15173 2025-08-12 cs.IR cs.AI cs.ET 57%

Recommendation with Generative Models

Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, Rene Vidal, Maheswaran Sathiamoorthy, Atoosa Kasrizadeh, Silvia Milano, Francesco Ricci

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments This submission is a full-length book, expanding significantly on two chapters previously submitted (arXiv:2409.10993v1, arXiv:2408.10946v1). It includes additional chapters, context, analysis, and content, providing a comprehensive presentation of the subject. We have ensured it is appropriately presented as a new, distinct work. arXiv admin note: substantial text overlap with arXiv:2409.10993

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04049 2025-08-07 cs.CV 57%

Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation

Jiayi He, Xu Wang, Shengeng Tang, Yaxiong Wang, Lechao Cheng, Dan Guo

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03690 2025-08-06 cs.CV cs.RO 57%

Veila: Panoramic LiDAR Generation from a Monocular RGB Image

Youquan Liu, Lingdong Kong, Weidong Yang, Ao Liang, Jianxiong Gao, Yang Wu, Xiang Xu, Xin Li, Linfeng Li, Runnan Chen, Ben Fei

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Preprint; 10 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03669 2025-08-06 cs.CV cs.RO 57%

OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World

Katherine Liu, Sergey Zakharov, Dian Chen, Takuya Ikeda, Greg Shakhnarovich, Adrien Gaidon, Rares Ambrus

机构 * Toyota Research Institute(丰田研究院) Woven by Toyota(丰田编织) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures. This version has typo fixes on top of the version published at ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03539 2025-08-06 cs.CV 57%

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

Long Qian, Bingke Zhu, Yingying Chen, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(自动化研究所基础模型研究中心,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Objecteye Inc.(Objecteye公司)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03535 2025-08-06 cs.CV 57%

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

Kaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu, Wenshuo Chen, Yutao Yue

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03320 2025-08-06 cs.CV 57%

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation

Peiyu Wang, Yi Peng, Yimeng Gan, Liang Hu, Tianyidan Xie, Xiaokun Wang, Yichen Wei, Chuanxin Tang, Bo Zhu, Changshi Li, Hongyang Wei, Eric Li, Xuchen Song, Yang Liu, Yahui Zhou

机构 * Multimodality Team, Skywork AI(Skywork AI 多模态团队)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00443 2025-08-05 cs.CV 57%

SDMatte: Grafting Diffusion Models for Interactive Matting

Longfei Huang, Yu Liang, Hao Zhang, Jinwei Chen, Wei Dong, Lunde Chen, Wanyu Liu, Bo Li, Peng-Tao Jiang

机构 * Shanghai University(上海大学) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted at ICCV 2025, 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17349 2025-08-05 cs.CV cs.IR 57%

DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition

Yiyan Xu, Wuqiang Zheng, Wenjie Wang, Fengbin Zhu, Xinting Hu, Yang Zhang, Fuli Feng, Tat-Seng Chua

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication in ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01778 2025-08-05 cs.CV cs.RO 57%

DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online HD Map Diffusion

Zhigang Sun, Yiru Wang, Anqing Jiang, Shuo Wang, Yu Gao, Yuwen Heng, Shouyi Zhang, An He, Hao Jiang, Jinhao Chai, Zichong Gu, Wang Jijun, Shichen Tang, Lavdim Halilaj, Juergen Luettin, Hao Sun

机构 * Bosch Corporate Research, Bosch (China) Investment Ltd.(博世企业研究院、博世(中国)投资有限公司) School of Communication and Information Engineering, Shanghai University(上海大学通信与信息工程学院) Shanghai Jiaotong University(上海交通大学) AIR, Tsinghua University(清华大学人工智能研究院) Robert Bosch GmbH(罗伯特·博世集团)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01761 2025-08-05 cs.LG cs.AI 57%

Semantically-Guided Inference for Conditional Diffusion Models: Enhancing Covariate Consistency in Time Series Forecasting

Rui Ding, Hanyang Meng, Zeyang Zhang, Jielong Yang

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏