arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2311.03054 2024-02-22 cs.CV 57%

AnyText: Multilingual Visual Text Generation And Editing

Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, Xuansong Xie

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03291 2024-02-22 cs.CV 57%

Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction

Yiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang, Soroush Vosoughi, Hongxia Yang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14603 2024-02-20 cs.CV 57%

Animate124: Animating One Image to 4D Dynamic Scene

Yuyang Zhao, Zhiwen Yan, Enze Xie, Lanqing Hong, Zhenguo Li, Gim Hee Lee

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Project Page: https://animate124.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09966 2024-02-16 cs.CV 57%

Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation

Junjie Shentu, Matthew Watson, Noura Al Moubayed

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08348 2024-02-14 cs.CV 57%

Visually Dehallucinative Instruction Generation

Sungguk Cha, Jusung Lee, Younghyun Lee, Cheoljong Yang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted in ICASSP2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12609 2024-02-13 cs.RO cs.AI cs.LG 57%

Denoising Heat-inspired Diffusion with Insulators for Collision Free Motion Planning

Junwoo Chang, Hyunwoo Ryu, Jiwoo Kim, Soochul Yoo, Jongeun Choi, Joohwan Seo, Nikhil Prakash, Roberto Horowitz

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments 9 pages, 6 figures

Journal ref NeurIPS 2023 Workshop on Diffusion Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06116 2024-02-12 cs.RO cs.AI 57%

LLMs for Coding and Robotics Education

Peng Shu, Huaqin Zhao, Hanqi Jiang, Yiwei Li, Shaochen Xu, Yi Pan, Zihao Wu, Zhengliang Liu, Guoyu Lu, Le Guan, Gong Chen, Xianqiao Wang Tianming Liu

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 20 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04717 2024-02-08 cs.CV 57%

InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior

Chenguo Lin, Yadong Mu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICLR 2024 for spotlight presentation; Project page: https://chenguolin.github.io/projects/InstructScene

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03754 2024-02-07 cs.CV 57%

Intensive Vision-guided Network for Radiology Report Generation

Fudan Zheng, Mengfei Li, Ying Wang, Weijiang Yu, Ruixuan Wang, Zhiguang Chen, Nong Xiao, Yutong Lu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by Physics in Medicine & Biology

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16762 2024-01-31 cs.CV 57%

Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization

Henglei Lv, Jiayu Xiao, Liang Li, Qingming Huang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12224 2024-01-24 cs.AR cs.AI 57%

LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation

Ruizhe Zhong, Xingbo Du, Shixiong Kai, Zhentao Tang, Siyuan Xu, Hui-Ling Zhen, Jianye Hao, Qiang Xu, Mingxuan Yuan, Junchi Yan

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09930 2024-01-24 cs.CV 57%

Deep Superpixel Generation and Clustering for Weakly Supervised Segmentation of Brain Tumors in MR Images

Jay J. Yoo, Khashayar Namdar, Farzad Khalvati

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 12 pages, LaTeX; updated methodology, added additional results, revised discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10934 2024-01-23 cs.IR cs.AI 57%

A New Creative Generation Pipeline for Click-Through Rate with Stable Diffusion Model

Hao Yang, Jianxin Yuan, Shuai Yang, Linhe Xu, Shuo Yuan, Yifan Zeng

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09985 2024-01-19 cs.CV 57%

WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Xiaofeng Wang, Zheng Zhu, Guan Huang, Boyuan Wang, Xinze Chen, Jiwen Lu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments project page: https://world-dreamer.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08332 2024-01-15 cs.CV 57%

Versatile Diffusion: Text, Images and Variations All in One Diffusion Model

Xingqian Xu, Zhangyang Wang, Eric Zhang, Kai Wang, Humphrey Shi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments ICCV 2023; Github link: https://github.com/SHI-Labs/Versatile-Diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10830 2024-01-11 cs.CV 57%

3D VR Sketch Guided 3D Shape Prototyping and Exploration

Ling Luo, Pinaki Nath Chowdhury, Tao Xiang, Yi-Zhe Song, Yulia Gryaditskaya

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02142 2024-01-09 cs.CV 57%

GUESS:GradUally Enriching SyntheSis for Text-Driven Human Motion Generation

Xuehao Gao, Yang Yang, Zhenyu Xie, Shaoyi Du, Zhongqian Sun, Yang Wu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Visualization and Computer Graphics (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01637 2024-01-04 cs.CL 57%

Social Media Ready Caption Generation for Brands

Himanshu Maheshwari, Koustava Goswami, Apoorv Saxena, Balaji Vasan Srinivasan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.10441 2024-01-01 cs.CV 57%

Synthetically Trained Icon Proposals for Parsing and Summarizing Infographics

Spandan Madan, Zoya Bylinskii, Matthew Tancik, Adrià Recasens, Kimberli Zhong, Sami Alsheikh, Hanspeter Pfister, Aude Oliva, Fredo Durand

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16141 2023-12-27 cs.CV 57%

VirtualPainting: Addressing Sparsity with Virtual Points and Distance-Aware Data Augmentation for 3D Object Detection

Sudip Dhakal, Dominic Carrillo, Deyuan Qu, Michael Nutt, Qing Yang, Song Fu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06718 2023-12-27 cs.AI 57%

Large Scale Foundation Models for Intelligent Manufacturing Applications: A Survey

Haotian Zhang, Semujju Stuart Dereck, Zhicheng Wang, Xianwei Lv, Kang Xu, Liang Wu, Ye Jia, Jing Wu, Zhuo Long, Wensheng Liang, X. G. Ma, Ruiyan Zhuang

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14733 2023-12-25 cs.CV 57%

Harnessing Diffusion Models for Visual Perception with Meta Prompts

Qiang Wan, Zilong Huang, Bingyi Kang, Jiashi Feng, Li Zhang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00087 2023-12-25 cs.CY cs.AI cs.HC 57%

Generative Artificial Intelligence in Learning Analytics: Contextualising Opportunities and Challenges through the Learning Analytics Cycle

Lixiang Yan, Roberto Martinez-Maldonado, Dragan Gašević

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11701 2023-12-20 eess.SY cs.CL cs.SY 57%

Opportunities and Challenges of Applying Large Language Models in Building Energy Efficiency and Decarbonization Studies: An Exploratory Overview

Liang Zhang, Zhelun Chen

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14197 2023-12-20 cs.CV 57%

PointVST: Self-Supervised Pre-training for 3D Point Clouds via View-Specific Point-to-Image Translation

Qijian Zhang, Junhui Hou

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted in IEEE TVCG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10674 2023-12-19 cs.CV 57%

A Framework of Full-Process Generation Design for Park Green Spaces Based on Remote Sensing Segmentation-GAN-Diffusion

Ran Chen, Xingjian Yi, Jing Zhao, Yueheng He, Bainian Chen, Xueqi Yao, Fangjun Liu, Haoran Li, Zeke Lian

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10631 2023-12-19 cs.NI cs.AI 57%

LLM-Twin: Mini-Giant Model-driven Beyond 5G Digital Twin Networking Framework with Semantic Secure Communication and Computation

Yang Hong, Jun Wu, Rosario Morello

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 18 pages, 11 figures, submitted to Scientific Reports on Nov. 12, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10682 2023-12-19 cs.CV cs.GR 57%

DiffStyler: Controllable Dual Diffusion for Text-Driven Image Stylization

Nisha Huang, Yuxin Zhang, Fan Tang, Chongyang Ma, Haibin Huang, Yong Zhang, Weiming Dong, Changsheng Xu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03667 2023-12-07 cs.CV 57%

WarpDiffusion: Efficient Diffusion Model for High-Fidelity Virtual Try-on

xujie zhang, Xiu Li, Michael Kampffmeyer, Xin Dong, Zhenyu Xie, Feida Zhu, Haoye Dong, Xiaodan Liang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03540 2023-12-07 cs.CV 57%

FoodFusion: A Latent Diffusion Model for Realistic Food Image Generation

Olivia Markham, Yuhao Chen, Chi-en Amy Tai, Alexander Wong

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏