arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2503.18589 2025-04-01 cs.CV 57%

Unified Uncertainty-Aware Diffusion for Multi-Agent Trajectory Modeling

Guillem Capellera, Antonio Rubio, Luis Ferraz, Antonio Agudo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15738 2025-04-01 cs.CV 57%

AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Qifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan, Yang Wu, Xiaoyang Wan, Juncheng Li, Siliang Tang, Hanwang Zhang, Yueting Zhuang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14239 2025-04-01 cs.CV 57%

MultiBooth: Towards Generating All Your Concepts in an Image from Text

Chenyang Zhu, Kai Li, Yue Ma, Chunming He, Xiu Li

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments To be published in AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22531 2025-03-31 eess.IV cs.CV cs.LG 57%

Deterministic Medical Image Translation via High-fidelity Brownian Bridges

Qisheng He, Nicholas Summerfield, Peiyong Wang, Carri Glide-Hurst, Ming Dong

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22079 2025-03-31 cs.CV 57%

A Semantic-Enhanced Heterogeneous Graph Learning Method for Flexible Objects Recognition

Kunshan Yang, Wenwei Luo, Yuguo Hu, Jiafu Yan, Mengmeng Jing, Lin Zuo

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21758 2025-03-28 cs.CV 57%

Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Qi Qin, Le Zhuo, Yi Xin, Ruoyi Du, Zhen Li, Bin Fu, Yiting Lu, Jiakang Yuan, Xinyue Li, Dongyang Liu, Xiangyang Zhu, Manyuan Zhang, Will Beddow, Erwann Millon, Victor Perez, Wenhai Wang, Conghui He, Bo Zhang, Xiaohong Liu, Hongsheng Li, Yu Qiao, Chang Xu, Peng Gao

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Tech Report, 21 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08503 2025-03-28 cs.CV 57%

StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

Mingkun Lei, Xue Song, Beier Zhu, Hao Wang, Chi Zhang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04314 2025-03-26 cs.CV 57%

Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

Zhanhao Liang, Yuhui Yuan, Shuyang Gu, Bohan Chen, Tiankai Hang, Mingxi Cheng, Ji Li, Liang Zheng

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments CVPR 2025. Project Page: https://rockeycoss.github.io/spo.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19312 2025-03-26 cs.CV 57%

ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning

Jiaqi Liao, Zhengyuan Yang, Linjie Li, Dianqi Li, Kevin Lin, Yu Cheng, Lijuan Wang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://ImageGen-CoT.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01110 2025-03-26 cs.CV 57%

RobustEMD: Domain Robust Matching for Cross-domain Few-shot Medical Image Segmentation

Yazhou Zhu, Minxian Li, Qiaolin Ye, Shidong Wang, Tong Xin, Haofeng Zhang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments More details should be included, and more experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07844 2025-03-26 cs.CV 57%

Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings

Arman Zarei, Keivan Rezaei, Samyadeep Basu, Mehrdad Saberi, Mazda Moayeri, Priyatham Kattakinda, Soheil Feizi

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18641 2025-03-25 cs.AI 57%

From Fragment to One Piece: A Survey on AI-Driven Graphic Design

Xingxing Zou, Wen Zhang, Nanxuan Zhao

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18461 2025-03-25 cs.CV 57%

MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Agentic Post-Processing

Lingting Zhu, Jingrui Ye, Runze Zhang, Zeyu Hu, Yingda Yin, Lanjiong Li, Jinnan Chen, Shengju Qian, Xin Wang, Qingmin Liao, Lequan Yu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 17 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18393 2025-03-25 cs.CV 57%

PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes

Xinhua Xu, Hong Liu, Jianbing Wu, Jinfu Liu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17784 2025-03-25 cs.AI 57%

MEPNet: Medical Entity-balanced Prompting Network for Brain CT Report Generation

Xiaodan Zhang, Yanzhao Shi, Junzhong Ji, Chengxin Zheng, Liangqiong Qu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments AAAI 2025 Oral Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15213 2025-03-25 cs.CV 57%

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

Qihao Liu, Xi Yin, Alan Yuille, Andrew Brown, Mannat Singh

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments CVPR 2025 camera-ready version. Project page: https://cross-flow.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16856 2025-03-25 cs.CV 57%

SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE

Yongwei Chen, Yushi Lan, Shangchen Zhou, Tengfei Wang, Xingang Pan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project page: https://cyw-3d.github.io/projects/SAR3D/ Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15959 2025-03-25 cs.RO cs.CV 57%

Diffusion Transformer Policy

Zhi Hou, Tianyi Zhang, Yuwen Xiong, Hengjun Pu, Chengyang Zhao, Ronglei Tong, Yu Qiao, Jifeng Dai, Yuntao Chen

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments preprint; New Project Page: https://robodita.github.io; revert unsuitable replacement

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17736 2025-03-25 eess.SP cs.CV cs.IT eess.IV math.IT 57%

Diffusion-Aided Joint Source Channel Coding For High Realism Wireless Image Transmission

Mingyu Yang, Bowen Liu, Boyang Wang, Hun-Seok Kim

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16978 2025-03-24 cs.AI 57%

Real-Time Diffusion Policies for Games: Enhancing Consistency Policies with Q-Ensembles

Ruoqi Zhang, Ziwei Luo, Jens Sjölund, Per Mattsson, Linus Gisslén, Alessandro Sestini

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16474 2025-03-24 cs.HC cs.AI 57%

From Voices to Worlds: Developing an AI-Powered Framework for 3D Object Generation in Augmented Reality

Majid Behravan, Denis Gracanin

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2502.15869

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07390 2025-03-24 cs.CV 57%

PersonaBooth: Personalized Text-to-Motion Generation

Boeun Kim, Hea In Jeong, JungHoon Sung, Yihua Cheng, Jeongmin Lee, Ju Yong Chang, Sang-Il Choi, Younggeun Choi, Saim Shin, Jungho Kim, Hyung Jin Chang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07685 2025-03-24 cs.CV 57%

Matrix3D: Large Photogrammetry Model All-in-One

Yuanxun Lu, Jingyang Zhang, Tian Fang, Jean-Daniel Nahmias, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao, Shiwei Li

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments CVPR 2025 camera ready. Project Page: https://nju-3dv.github.io/projects/matrix3d

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16153 2025-03-21 cs.CV 57%

FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing

Tianyi Wei, Yifan Zhou, Dongdong Chen, Xingang Pan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project page: https://wtybest.github.io/projects/FreeFlux/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02210 2025-03-19 cs.CV 57%

One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering

Deepayan Das, Davide Talon, Massimiliano Mancini, Yiming Wang, Elisa Ricci

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06350 2025-03-19 cs.CV 57%

T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Kaiyi Huang, Chengqi Duan, Kaiyue Sun, Enze Xie, Zhenguo Li, Xihui Liu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments This is the journal version. For conference version (T2I-CompBench): arXiv:2307.06350v2. Project page: https://karine-h.github.io/T2I-CompBench-new/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13436 2025-03-18 cs.CV cs.LG 57%

Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

Lijie Fan, Luming Tang, Siyang Qin, Tianhong Li, Xuan Yang, Siyuan Qiao, Andreas Steiner, Chen Sun, Yuanzhen Li, Tao Zhu, Michael Rubinstein, Michalis Raptis, Deqing Sun, Radu Soricut

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12472 2025-03-18 cs.CV 57%

Diffusion-based Synthetic Data Generation for Visible-Infrared Person Re-Identification

Wenbo Dai, Lijing Lu, Zhihang Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12466 2025-03-18 cs.RO cs.CV 57%

Modality-Composable Diffusion Policy via Inference-Time Distribution-level Composition

Jiahang Cao, Qiang Zhang, Hanzhong Guo, Jiaxu Wang, Hao Cheng, Renjing Xu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICLR 2025 Generative Models for Robot Learning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12383 2025-03-18 cs.CV 57%

VRsketch2Gaussian: 3D VR Sketch Guided 3D Object Generation with Gaussian Splatting

Songen Gu, Haoxuan Song, Binjie Liu, Qian Yu, Sanyi Zhang, Haiyong Jiang, Jin Huang, Feng Tian

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏