arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2411.16749 2024-12-03 cs.CV 57%

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks

You Li, Fan Ma, Yi Yang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18301 2024-11-28 cs.CV 57%

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Tianyi Wei, Dongdong Chen, Yifan Zhou, Xingang Pan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00345 2024-11-28 cs.CV 57%

MaGRITTe: Manipulative and Generative 3D Realization from Image, Topview and Text

Takayuki Hara, Tatsuya Harada

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://hara012.github.io/MaGRITTe-project

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17673 2024-11-27 cs.CV 57%

SketchAgent: Language-Driven Sequential Sketch Generation

Yael Vinker, Tamar Rott Shaham, Kristine Zheng, Alex Zhao, Judith E Fan, Antonio Torralba

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments project page: https://sketch-agent.csail.mit.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17423 2024-11-27 cs.CV 57%

DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters

Mingze Sun, Junhao Chen, Junting Dong, Yurun Chen, Xinyu Jiang, Shiwei Mao, Puhua Jiang, Jingbo Wang, Bo Dai, Ruqi Huang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17221 2024-11-27 cs.CV 57%

AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM

Jiarui Wang, Huiyu Duan, Guangtao Zhai, Juntong Wang, Xiongkuo Min

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14957 2024-11-26 cs.CL 57%

Information Extraction from Heterogeneous Documents without Ground Truth Labels using Synthetic Label Generation and Knowledge Distillation

Aniket Bhattacharyya, Anurag Tripathi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments Accepted to WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15236 2024-11-26 cs.CV cs.LG 57%

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

Jeeyung Kim, Erfan Esmaeili, Qiang Qiu

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15034 2024-11-25 cs.CV cs.LG 57%

HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads

Yu Xu, Fan Tang, Juan Cao, Yuxin Zhang, Xiaoyu Kong, Jintao Li, Oliver Deussen, Tong-Yee Lee

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03181 2024-11-20 cs.LG cs.AI cs.RO 57%

Behavior Generation with Latent Actions

Seungjae Lee, Yibin Wang, Haritheja Etukuru, H. Jin Kim, Nur Muhammad Mahi Shafiullah, Lerrel Pinto

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Github repo: https://github.com/jayLEE0301/vq_bet_official

Journal ref PMLR 235:26991-27008, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18530 2024-11-19 cs.CV 57%

MatchTime: Towards Automatic Soccer Game Commentary Generation

Jiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang, Weidi Xie

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by EMNLP 2024 (Oral Presentation); Project Page: https://haoningwu3639.github.io/MatchTime/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08196 2024-11-14 cs.CV 57%

Latent Space Disentanglement in Diffusion Transformers Enables Precise Zero-shot Semantic Editing

Zitao Shuai, Chenwei Wu, Zhengxu Tang, Bowen Song, Liyue Shen

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2408.13335

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01956 2024-11-13 cs.CV 57%

Enhance Image-to-Image Generation with LLaVA-generated Prompts

Zhicheng Ding, Panfeng Li, Qikai Yang, Siyang Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted by 2024 5th International Conference on Information Science, Parallel and Distributed Systems

Journal ref Proceedings of the 2024 5th International Conference on Information Science, Parallel and Distributed Systems (ISPDS), 2024, pp. 77-81

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02504 2024-11-13 eess.IV cs.CV 57%

Functional Imaging Constrained Diffusion for Brain PET Synthesis from Structural MRI

Minhui Yu, Mengqi Wu, Ling Yue, Andrea Bozoki, Mingxia Liu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01698 2024-11-12 eess.IV cs.CV 57%

COSMIC: Compress Satellite Images Efficiently via Diffusion Compensation

Ziyuan Zhang, Han Qiu, Maosen Zhang, Jun Liu, Bin Chen, Tianwei Zhang, Hewu Li

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09774 2024-11-07 cs.CV 57%

Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences through $f$-divergence Minimization

Haoyuan Sun, Bo Xia, Yongzhe Chang, Xueqian Wang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14138 2024-11-05 cs.CV 57%

Visual Text Generation in the Wild

Yuanzhi Zhu, Jiawei Liu, Feiyu Gao, Wenyu Liu, Xinggang Wang, Peng Wang, Fei Huang, Cong Yao, Zhibo Yang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14530 2024-11-05 cs.CV 57%

Multistable Shape from Shading Emerges from Patch Diffusion

Xinran Nicole Han, Todd Zickler, Ko Nishino

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12470 2024-11-04 cs.CV eess.IV 57%

HSIGene: A Foundation Model For Hyperspectral Image Generation

Li Pang, Xiangyong Cao, Datao Tang, Shuang Xu, Xueru Bai, Feng Zhou, Deyu Meng

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23803 2024-11-01 cs.HC cs.AI 57%

Generative AI for Accessible and Inclusive Extended Reality

Jens Grubert, Junlong Chen, Per Ola Kristensson

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Presented at the CHI 2024 Workshop "Building a Metaverse for All: Opportunities and Challenges for Future Inclusive and Accessible Virtual Environments", May 11, 2024, Honolulu, Hawaii

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23676 2024-11-01 cs.CV 57%

Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach

Mathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet Iscen

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22149 2024-10-30 cs.CV 57%

Capacity Control is an Effective Memorization Mitigation Mechanism in Text-Conditional Diffusion Models

Raman Dutt, Pedro Sanchez, Ondrej Bohdal, Sotirios A. Tsaftaris, Timothy Hospedales

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted at the GenLaw (Generative AI + Law) workshop at ICML'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19079 2024-10-30 cs.CV cs.LG 57%

BIFRÖST: 3D-Aware Image compositing with Language Instructions

Lingxiao Li, Kaixiong Gong, Weihong Li, Xili Dai, Tao Chen, Xiaojun Yuan, Xiangyu Yue

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

Comments NeurIPS 2024, Code Available: https://github.com/lingxiao-li/Bifrost

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20164 2024-10-29 cs.LG cs.CV 57%

Prompt Diffusion Robustifies Any-Modality Prompt Learning

Yingjun Du, Gaowen Liu, Yuzhang Shang, Yuguang Yao, Ramana Kompella, Cees G. M. Snoek

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19235 2024-10-28 cs.RO cs.AI 57%

Learning Diffusion Policies from Demonstrations For Compliant Contact-rich Manipulation

Malek Aburub, Cristian C. Beltran-Hernandez, Tatsuya Kamijo, Masashi Hamaya

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00505 2024-10-24 cs.CV 57%

Improving Text Generation on Images with Synthetic Captions

Jun Young Koh, Sang Hyun Park, Joy Song

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 2024 16th IIAI International Congress on Advanced Applied Informatics (IIAI-AAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16840 2024-10-23 cs.CV 57%

MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model

Meng Xu, Tong Zhang, Fuyun Wang, Yi Lei, Xin Liu, Zhen Cui

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16820 2024-10-23 cs.CV 57%

AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models

Yongjian Wu, Yang Zhou, Jiya Saiyin, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yan Xu

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments This article has been accepted for publication in a future issue of IEEE Transactions on Medical Imaging (TMI), but has not been fully edited. Content may change prior to final publication. Citation information: DOI: https://doi.org/10.1109/TMI.2024.3473745 . Code: https://github.com/wuyongjianCODE/AttriPrompter

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16395 2024-10-23 cs.CV cs.GR 57%

Joker: Conditional 3D Head Synthesis with Extreme Facial Expressions

Malte Prinzler, Egor Zakharov, Vanessa Sklyarova, Berna Kabadayi, Justus Thies

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Project Page: https://malteprinzler.github.io/projects/joker/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11965 2024-10-23 cs.CV 57%

UrbanWorld: An Urban World Model for 3D City Generation

Yu Shang, Yuming Lin, Yu Zheng, Hangyu Fan, Jingtao Ding, Jie Feng, Jiansheng Chen, Li Tian, Yong Li

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏