arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2402.01293 2024-07-23 cs.LG cs.CL 57%

Can MLLMs Perform Text-to-Image In-Context Learning?

Yuchen Zeng, Wonjun Kang, Yicong Chen, Hyung Il Koo, Kangwook Lee

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments Accepted at COLM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13181 2024-07-19 cs.CV 57%

Training-Free Large Model Priors for Multiple-in-One Image Restoration

Xuanhua He, Lang Li, Yingying Wang, Hui Zheng, Ke Cao, Keyu Yan, Rui Li, Chengjun Xie, Jie Zhang, Man Zhou

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18958 2024-07-19 cs.CV 57%

AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation

Yanan Sun, Yanchen Liu, Yinhao Tang, Wenjie Pei, Kai Chen

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ECCV 2024, code and dataset available in https://github.com/open-mmlab/AnyControl

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12727 2024-07-18 cs.CV 57%

NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model

Zhongqun Zhang, Hengfei Wang, Ziwei Yu, Yihua Cheng, Angela Yao, Hyung Jin Chang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10179 2024-07-18 cs.CV 57%

Animate Your Motion: Turning Still Images into Dynamic Videos

Mingxiao Li, Bo Wan, Marie-Francine Moens, Tinne Tuytelaars

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted at European Conference on Computer Vision (ECCV 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15657 2024-07-18 cs.CV 57%

Enhancing Diffusion Models with Text-Encoder Reinforcement Learning

Chaofeng Chen, Annan Wang, Haoning Wu, Liang Liao, Wenxiu Sun, Qiong Yan, Weisi Lin

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10937 2024-07-16 cs.CV 57%

IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation

Yuanhao Zhai, Kevin Lin, Linjie Li, Chung-Ching Lin, Jianfeng Wang, Zhengyuan Yang, David Doermann, Junsong Yuan, Zicheng Liu, Lijuan Wang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments ECCV 2024; project page: https://yhzhai.github.io/idol/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07516 2024-07-15 cs.CV 57%

Instant 3D Human Avatar Generation using Image Diffusion Models

Nikos Kolotouros, Thiemo Alldieck, Enric Corona, Eduard Gabriel Bazavan, Cristian Sminchisescu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03961 2024-07-12 cs.CV cs.LG 57%

Leveraging Latent Diffusion Models for Training-Free In-Distribution Data Augmentation for Surface Defect Detection

Federico Girella, Ziyue Liu, Franco Fummi, Francesco Setti, Marco Cristani, Luigi Capogrosso

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted at the 21st International Conference on Content-Based Multimedia Indexing (CBMI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16075 2024-07-12 cs.LG cs.AI cs.RO 57%

Don't Start from Scratch: Behavioral Refinement via Interpolant-based Policy Diffusion

Kaiqi Chen, Eugene Lim, Kelvin Lin, Yiyang Chen, Harold Soh

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14473 2024-07-12 eess.IV cs.CV 57%

Joint Diffusion: Mutual Consistency-Driven Diffusion Model for PET-MRI Co-Reconstruction

Taofeng Xie, Zhuo-Xu Cui, Chen Luo, Huayu Wang, Congcong Liu, Yuanzhi Zhang, Xuemei Wang, Yanjie Zhu, Guoqing Chen, Dong Liang, Qiyu Jin, Yihang Zhou, Haifeng Wang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05323 2024-07-09 eess.IV cs.CV 57%

Enhancing Label-efficient Medical Image Segmentation with Text-guided Diffusion Models

Chun-Mei Feng

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments MICCAI 2024, Early Accept

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10842 2024-07-04 cs.CV 57%

Automated Radiology Report Generation: A Review of Recent Advances

Phillip Sloan, Philip Clatworthy, Edwin Simpson, Majid Mirmehdi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 24 pages, 8 figures, 6 tables. Accepted by IEEE Reviews in Biomedical Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01624 2024-07-03 cs.LG cs.AI 57%

Guided Trajectory Generation with Diffusion Models for Offline Model-based Optimization

Taeyoung Yun, Sujin Yun, Jaewoo Lee, Jinkyoo Park

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments 29 pages, 11 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01104 2024-07-02 cs.CV 57%

Semantic-guided Adversarial Diffusion Model for Self-supervised Shadow Removal

Ziqi Zeng, Chen Zhao, Weiling Cai, Chenyu Dong

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20469 2024-07-02 cs.CV 57%

Is Synthetic Data all We Need? Benchmarking the Robustness of Models Trained with Synthetic Images

Krishnakant Singh, Thanush Navaratnam, Jannik Holmer, Simone Schaub-Meyer, Stefan Roth

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted at CVPR 2024 Workshop: SyntaGen-Harnessing Generative Models for Synthetic Visual Datasets. Project page at https://synbenchmark.github.io/SynCloneBenchmark Comments: Fix typo in Fig. 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14436 2024-06-21 cs.CV cs.RO 57%

Video Generation with Learned Action Prior

Meenakshi Sarkar, Devansh Bhardwaj, Debasish Ghose

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12030 2024-06-18 cs.CL 57%

Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation

Masato Mita, Soichiro Murakami, Akihiko Kato, Peinan Zhang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments Accepted to ACL2024 (main, long)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05945 2024-06-14 cs.CV 57%

Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Peng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du, Xu Luo, Longtian Qiu, Yuhang Zhang, Chen Lin, Rongjie Huang, Shijie Geng, Renrui Zhang, Junlin Xi, Wenqi Shao, Zhengkai Jiang, Tianshuo Yang, Weicai Ye, He Tong, Jingwen He, Yu Qiao, Hongsheng Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Technical Report; Code at: https://github.com/Alpha-VLLM/Lumina-T2X

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08392 2024-06-13 cs.CV 57%

FontStudio: Shape-Adaptive Diffusion Model for Coherent and Consistent Font Effect Generation

Xinzhi Mu, Li Chen, Bohan Chen, Shuyang Gu, Jianmin Bao, Dong Chen, Ji Li, Yuhui Yuan

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Project-page: https://font-studio.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05565 2024-06-11 cs.CV 57%

Medical Vision Generalist: Unifying Medical Imaging Tasks in Context

Sucheng Ren, Xiaoke Huang, Xianhang Li, Junfei Xiao, Jieru Mei, Zeyu Wang, Alan Yuille, Yuyin Zhou

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13125 2024-06-11 eess.IV cs.CV 57%

Deep Learning Approaches for Data Augmentation in Medical Imaging: A Review

Aghiles Kebaili, Jérôme Lapuyade-Lahorgue, Su Ruan

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04712 2024-06-10 cs.CL 57%

AICoderEval: Improving AI Domain Code Generation of Large Language Models

Yinghui Xia, Yuyan Chen, Tianyu Shi, Jun Wang, Jinsong Yang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03688 2024-06-07 eess.IV cs.CV 57%

Shadow and Light: Digitally Reconstructed Radiographs for Disease Classification

Benjamin Hou, Qingqing Zhu, Tejas Sudarshan Mathai, Qiao Jin, Zhiyong Lu, Ronald M. Summers

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16835 2024-06-06 cs.CV 57%

Unified-modal Salient Object Detection via Adaptive Prompt Learning

Kunpeng Wang, Chenglong Li, Zhengzheng Tu, Zhengyi Liu, Bin Luo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 13 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01489 2024-06-05 cs.CV 57%

DA-HFNet: Progressive Fine-Grained Forgery Image Detection and Localization Based on Dual Attention

Yang Liu, Xiaofei Li, Jun Zhang, Shengze Hu, Jun Lei

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00644 2024-06-04 cs.CV 57%

Ultrasound Report Generation with Cross-Modality Feature Alignment via Unsupervised Guidance

Jun Li, Tongkun Su, Baoliang Zhao, Faqin Lv, Qiong Wang, Nassir Navab, Ying Hu, Zhongliang Jiang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09808 2024-06-04 cs.CL 57%

PixT3: Pixel-based Table-To-Text Generation

Iñigo Alonso, Eneko Agirre, Mirella Lapata

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00121 2024-06-04 cs.CV 57%

Empowering Visual Creativity: A Vision-Language Assistant to Image Editing Recommendations

Tiancheng Shen, Jun Hao Liew, Long Mai, Lu Qi, Jiashi Feng, Jiaya Jia

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18700 2024-05-31 cs.CV 57%

Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion Prediction

Xuehao Gao, Yang Yang, Yang Wu, Shaoyi Du, Guo-Jun Qi

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏