arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2505.08817 2025-05-15 cs.CV cs.LG 57%

Towards SFW sampling for diffusion models via external conditioning

Camilo Carvajal Reyes, Joaquín Fontbona, Felipe Tobar

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepcted at IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18715 2025-05-15 cs.CV 57%

BCTR: Bidirectional Conditioning Transformer for Scene Graph Generation

Peng Hao, Weilong Wang, Xiaobing Wang, Yingying Jiang, Hanchao Jia, Shaowei Cui, Junhang Wei, Xiaoshuai Hao

机构 * Samsung R&D Institute China–Beijing(三星中国北京研发中心) School of Mathematics, Southwestern University of Finance and Economics(西南财经大学数学学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能实验室) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08281 2025-05-14 cs.CV eess.IV 57%

Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion

Anle Ke, Xu Zhang, Tong Chen, Ming Lu, Chao Zhou, Jiawen Gu, Zhan Ma

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07721 2025-05-13 cs.CV 57%

Gameplay Highlights Generation

Vignesh Edithal, Le Zhang, Ilia Blank, Imran Junejo

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07057 2025-05-13 cs.CV 57%

DAPE: Dual-Stage Parameter-Efficient Fine-Tuning for Consistent Video Editing with Diffusion Models

Junhao Xia, Chaoyang Zhang, Yecheng Zhang, Chengyang Zhou, Zhichang Wang, Bochun Liu, Dongshuo Yin

机构 * Tsinghua University(清华大学) Duke University(杜克大学) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11735 2025-05-12 cs.AI 57%

Clinical Insights: A Comprehensive Review of Language Models in Medicine

Nikita Neveditsin, Pawan Lingras, Vijay Mago

机构 * Department of Mathematics and Computing Science, Saint Mary’s University(数学与计算科学系,圣玛丽大学) School of Health Policy and Management, York University(健康政策与管理学院,约克大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Submitted to PLOS Digital Health, Revision 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05516 2025-05-12 q-bio.TO cs.AI cs.HC 57%

AI-powered virtual eye: perspective, challenges and opportunities

Yue Wu, Yibo Guo, Yulong Yan, Jiancheng Yang, Xin Zhou, Ching-Yu Cheng, Danli Shi, Mingguang He

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 30 Pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07085 2025-05-12 cs.RO cs.CV 57%

RS2AD: End-to-End Autonomous Driving Data Generation from Roadside Sensor Observations

Ruidan Xing, Runyi Huang, Qing Xu, Lei He

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院) State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(清华大学智能绿色车辆与移动系统国家重点实验室) School of Instrumentation and Optoelectronic Engineering, BeiHang University(北航仪器与光电工程学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05367 2025-05-09 cs.CV eess.IV 57%

Joint Super-Resolution and Segmentation for 1-m Impervious Surface Area Mapping in China's Yangtze River Economic Belt

Jie Deng, Danfeng Hong, Chenyu Li, Naoto Yokoya

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China(中国科学院 aerospace information research institute, 北京) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100049, China(中国科学院大学电子电气与通信工程学院, 北京) School of Mathematics, Southeast University, Nanjing 210096, China(东南大学数学学院, 南京) Graduate School of Frontier Sciences, the University of Tokyo, Chiba 277-8561, Japan(东京大学前沿科学研究生院, 日本) RIKEN Center for Advanced Intelligence Project, Tokyo 103-0027, Japan(RIKEN 高度智能项目中心, 日本)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02417 2025-05-09 cs.LG cs.AI 57%

T2S: High-resolution Time Series Generation with Text-to-Series Diffusion Models

Yunfeng Ge, Jiawei Li, Yiji Zhao, Haomin Wen, Zhao Li, Meikang Qiu, Hongyan Li, Ming Jin, Shirui Pan

机构 * Xidian University(西安电子科技大学) Griffith University(格里菲斯大学) Yunnan University(云南大学) Carnegie Mellon University(卡内基梅隆大学) Zhejiang University(浙江大学) Augusta University(奥古斯塔大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Accepted by the 34th International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03072 2025-05-08 cs.RO cs.AI cs.MA 57%

Multi-Robot Motion Planning with Diffusion Models

Yorai Shaoul, Itamar Mishani, Shivam Vats, Jiaoyang Li, Maxim Likhachev

机构 * Carnegie Mellon University(卡内基梅隆大学) Brown University(布朗大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments The first three authors contributed equally to this work. Published at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03217 2025-05-07 cs.NE cs.AI 57%

Accelerating Evolution: Integrating PSO Principles into Real-Coded Genetic Algorithm Crossover

Xiaobo Jin, JiaShu Tu

机构 * Xiaobo Jin School of Information Engineering Taizhou Vocational College of Science & Technology(金晓波 学校信息工程学院 太zhou 职业科技学院) Traditional Chinese Medicine department Taizhou First People’s Hospital(传统中医部门 太zhou 第一人民医院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 14 pages,2 figures,4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02329 2025-05-06 cs.CV 57%

MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis

Dewei Zhou, You Li, Fan Ma, Zongxin Yang, Yi Yang

机构 * ReLER, CCAI, Zhejiang University(ReLER、CCAI、浙江大学) DBMI, HMS, Harvard University(DBMI、HMS、哈佛大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21814 2025-05-01 cs.CV 57%

Why Compress What You Can Generate? When GPT-4o Generation Ushers in Image Compression Fields

Yixin Gao, Xiaohan Pan, Xin Li, Zhibo Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21423 2025-05-01 cs.CV 57%

Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision

Weicai Yan, Wang Lin, Zirun Guo, Ye Wang, Fangming Feng, Xiaoda Yang, Zehan Wang, Tao Jin

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21334 2025-05-01 cs.CV 57%

Simple Visual Artifact Detection in Sora-Generated Videos

Misora Sugiyama, Hirokatsu Kataoka

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19614 2025-04-29 cs.CV 57%

DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer

Junpeng Jiang, Gangyi Hong, Miao Zhang, Hengtong Hu, Kun Zhan, Rui Shao, Liqiang Nie

机构 * Harbin Institute of Technology(哈尔滨工业大学) Li Auto Inc.(Li汽车公司) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17789 2025-04-29 cs.CV 57%

Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models

Xu Ma, Peize Sun, Haoyu Ma, Hao Tang, Chih-Yao Ma, Jialiang Wang, Kunpeng Li, Xiaoliang Dai, Yujun Shi, Xuan Ju, Yushi Hu, Artsiom Sanakoyeu, Felix Juefei-Xu, Ji Hou, Junjiao Tian, Tao Xu, Tingbo Hou, Yen-Cheng Liu, Zecheng He, Zijian He, Matt Feiszli, Peizhao Zhang, Peter Vajda, Sam Tsai, Yun Fu

机构 * Northeastern University(东北大学) Meta GenAI(Meta 生成人工智能) Meta FAIR National University of Singapore(国立新加坡大学) The Chinese University of Hong Kong(香港中文大学) University of Washington(华盛顿大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://ma-xu.github.io/token-shuffle/ Add related works

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23452 2025-04-29 cs.CV 57%

VideoGen-Eval: Agent-based System for Video Generation Evaluation

Yuhang Yang, Ke Fan, Shangkun Sun, Hongxiang Li, Ailing Zeng, FeiLin Han, Wei Zhai, Wei Liu, Yang Cao, Zheng-Jun Zha

机构 * USTC(中国科学技术大学) SJTU(上海交通大学) PKUSZ(北京大学软件学院) Tencent(腾讯) BFA(北京航空航天大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

Comments project:https://github.com/AILab-CVC/VideoGen-Eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17267 2025-04-25 cs.HC cs.MM 57%

MV-Crafter: An Intelligent System for Music-guided Video Generation

Chuer Chen, Shengqi Dang, Yuqi Liu, Nanxuan Zhao, Yang Shi, Nan Cao

专题命中 多模态生成 :audio-visual(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12652 2025-04-24 cs.CV 57%

UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

Tsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu, Zhe Gan, Yinfei Yang

机构 * Apple(苹果公司)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04105 2025-04-24 cs.AI 57%

Natural Language Processing in the Patent Domain: A Survey

Lekang Jiang, Stephan Goetz

机构 * University of Cambridge(剑桥大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Published in Artificial Intelligence Review

Journal ref Artif Intell Rev 58, 214 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16080 2025-04-23 cs.CV 57%

From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning

Le Zhuo, Liangbing Zhao, Sayak Paul, Yue Liao, Renrui Zhang, Yi Xin, Peng Gao, Mohamed Elhoseiny, Hongsheng Li

机构 * CUHK MMLab(香港中文大学多模态实验室) KAUST(科威特科学与技术研究中心) Hugging Face(Hugging Face公司) Shanghai AI Lab(上海人工智能实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments All code, checkpoints, and datasets are available at \url{https://diffusion-cot.github.io/reflection2perfection}

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15009 2025-04-22 cs.CV 57%

Insert Anything: Image Insertion via In-Context Editing in DiT

Wensong Song, Hong Jiang, Zongxing Yang, Ruijie Quan, Yi Yang

机构 * Zhejiang University(浙江大学) Harvard University(哈佛大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10329 2025-04-22 cs.CV 57%

InstructEngine: Instruction-driven Text-to-Image Alignment

Xingyu Lu, Yuhang Hu, YiFan Zhang, Kaiyu Jiang, Changyi Liu, Tianke Zhang, Jinpeng Wang, Chun Yuan, Bin Wen, Fan Yang, Tingting Gao, Di Zhang

机构 * Tsinghua University(清华大学) Kuaishou Technology(快手科技) CASIA(中国科学院自动化研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10148 2025-04-22 cs.CV 57%

Hierarchical and Step-Layer-Wise Tuning of Attention Specialty for Multi-Instance Synthesis in Diffusion Transformers

Chunyang Zhang, Zhenhong Sun, Zhicheng Zhang, Junyan Wang, Yu Zhang, Dong Gong, Huadong Mo, Daoyi Dong

机构 * School of Systems and Computing University of New South Wales(系统与计算学院 新南威尔士大学) School of Engineering Australian National University(工程学院 澳大利亚国立大学) School of Business University of New South Wales(商学院 新南威尔士大学) Australian Institute for Machine Learning University of Adelaide(机器学习研究所 阿德莱德大学) School of Computer Science and Engineering University of New South Wales(计算机科学与工程学院 新南威尔士大学) School of Computer Science University of Technology Sydney(计算机科学学院 技术大学悉尼)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11603 2025-04-21 cs.CV 57%

LaMD: Latent Motion Diffusion for Image-Conditional Video Generation

Yaosi Hu, Zhenzhong Chen, Chong Luo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments accepted by IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13162 2025-04-18 cs.CV 57%

Personalized Text-to-Image Generation with Auto-Regressive Models

Kaiyue Sun, Xian Liu, Yao Teng, Xihui Liu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project page: https://github.com/KaiyueSun98/T2I-Personalization-with-AR

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12540 2025-04-18 cs.GR cs.CV cs.RO 57%

UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character Control

Yan Wu, Korrawe Karunratanakul, Zhengyi Luo, Siyu Tang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Project page: https://wuyan01.github.io/uniphys-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11825 2025-04-17 eess.IV cs.CV 57%

TextDiffSeg: Text-guided Latent Diffusion Model for 3d Medical Images Segmentation

Kangbo Ma

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏