arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2506.02485 2025-08-05 cs.AI cs.CE 57%

Generative AI as a Pillar for Predicting 2D and 3D Wildfire Spread: Beyond Physics-Based Models and Traditional Deep Learning

Haowen Xu, Sisi Zlatanova, Ruiyu Liang, Ismet Canbulat

机构 * GRID, School of Built Environment, UNSW Sydney, NSW 2052 Australia(GRID,环境建筑学院,新南威尔士大学悉尼分校) School of Minerals and Energy Resources Engineering, UNSW Sydney, NSW 2052 Australia(矿物与能源资源工程学院,新南威尔士大学悉尼分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11935 2025-08-05 cs.CL 57%

ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing

Xuanle Zhao, Xuexin Liu, Haoyue Yang, Xianzhen Luo, Fanhu Zeng, Jianling Li, Qi Shi, Chi Chen

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Harbin Institute of Technology(哈尔滨工业大学) Tianjin University(天津大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments Accepted by ACL2025 Findings, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13490 2025-08-05 cs.CV 57%

Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing

Joowon Kim, Ziseok Lee, Donghyeon Cho, Sanghyun Jo, Yeonsung Jung, Kyungsu Kim, Eunho Yang

机构 * KAIST(韩国科学技术院) Seoul National University(首尔国立大学) OGQ

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07580 2025-08-05 cs.CV 57%

InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior

Chenguo Lin, Yuchen Lin, Panwang Pan, Xuanyang Zhang, Yadong Mu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) PICO AI group, ByteDance(字节跳动PICO AI团队)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to T-PAMI 2025. This paper is an extension of ICLR 2024 "InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior". arXiv admin note: substantial text overlap with arXiv:2402.04717

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01130 2025-08-05 cs.CV 57%

Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models

Bicheng Xu, Qi Yan, Renjie Liao, Lele Wang, Leonid Sigal

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18004 2025-08-04 cs.AI 57%

E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI

Yusen Peng, Shuhua Mao

机构 * University of Warwick(沃里克大学) Wuhan University of Technology(武汉理工大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI

Comments 44 pages,11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23178 2025-08-01 cs.SE cs.AI 57%

AutoBridge: Automating Smart Device Integration with Centralized Platform

Siyuan Liu, Zhice Yang, Huangxun Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ShanghaiTech University(上海科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 14 pages, 12 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22789 2025-08-01 cs.LG cs.AI 57%

G-Core: A Simple, Scalable and Balanced RLHF Trainer

Junyu Wu, Weiming Chang, Xiaotao Liu, Guanyou He, Haoqiang Hong, Boqi Liu, Hongtao Tian, Tao Yang, Yunsheng Shi, Feng Lin, Ting Yao

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments I haven't received company approval yet, and I uploaded it by mistake

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17761 2025-08-01 cs.CV 57%

Step1X-Edit: A Practical Framework for General Image Editing

Shiyu Liu, Yucheng Han, Peng Xing, Fukun Yin, Rui Wang, Wei Cheng, Jiaqi Liao, Yingming Wang, Honghao Fu, Chunrui Han, Guopeng Li, Yuang Peng, Quan Sun, Jingwei Wu, Yan Cai, Zheng Ge, Ranchen Ming, Lei Xia, Xianfang Zeng, Yibo Zhu, Binxing Jiao, Xiangyu Zhang, Gang Yu, Daxin Jiang

机构 * Step1X-Image Team(Step1X-图像团队)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments code: https://github.com/stepfun-ai/Step1X-Edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05422 2025-08-01 cs.CV cs.LG cs.RO 57%

EP-Diffuser: An Efficient Diffusion Model for Traffic Scene Generation and Prediction via Polynomial Representations

Yue Yao, Mohamed-Khalil Bouzidi, Daniel Goehring, Joerg Reichardt

机构 * Department of Mathematics and Computer Science, Freie Universität Berlin(数学与计算机科学系,柏林自由大学) Continental Automotive GmbH(大陆汽车股份有限公司)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21627 2025-07-30 cs.CV 57%

GuidPaint: Class-Guided Image Inpainting with Diffusion Models

Qimin Wang, Xinda Liu, Guohua Geng

机构 * Northwest University(西北大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20976 2025-07-29 cs.CV 57%

Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision

Xiao Fang, Minhyek Jeon, Zheyang Qin, Stanislav Panev, Celso de Melo, Shuowen Hu, Shayok Chakraborty, Fernando De la Torre

机构 * Carnegie Mellon University(卡内基梅隆大学) DEVCOM Army Research Laboratory(陆军研究实验室) Florida State University(佛罗里达州立大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19882 2025-07-29 cs.AI 57%

Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation

Xinshu Li, Ruoyu Wang, Erdun Gao, Mingming Gong, Lina Yao

机构 * The University of New South Wales(新南威尔士大学) The University of Adelaide(阿德莱德大学) The University of Melbourne(墨尔本大学) CSIRO’s Data 61(CSIRO数据61)

专题命中 多模态生成 :image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21771 2025-07-29 cs.CV 57%

A Unified Image-Dense Annotation Generation Model for Underwater Scenes

Hongkai Lin, Dingkang Liang, Zhenghao Qi, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025. The code is available at https://github.com/HongkLin/TIDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11824 2025-07-29 cs.CV 57%

KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

Hsin-Ping Huang, Xinyi Wang, Yonatan Bitton, Hagai Taitelbaum, Gaurav Singh Tomar, Ming-Wei Chang, Xuhui Jia, Kelvin C. K. Chan, Hexiang Hu, Yu-Chuan Su, Ming-Hsuan Yang

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

Comments Project page: https://kitten-project.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17801 2025-07-25 cs.CV 57%

Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling

Yi Xin, Juncheng Yan, Qi Qin, Zhen Li, Dongyang Liu, Shicheng Li, Victor Shea-Jay Huang, Yupeng Zhou, Renrui Zhang, Le Zhuo, Tiancheng Han, Xiaoqing Sun, Siqi Luo, Mengmeng Wang, Bin Fu, Yuewen Cao, Hongsheng Li, Guangtao Zhai, Xiaohong Liu, Yu Qiao, Peng Gao

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Nanjing University(南京大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University of Technology(浙江工业大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Tech Report, 23 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17774 2025-07-25 cs.HC cs.AI 57%

Human-AI Co-Creation: A Framework for Collaborative Design in Intelligent Systems

Zhangqi Liu

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12890 2025-07-25 eess.AS cs.SD 57%

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

Huakang Chen, Yuepeng Jiang, Guobin Ma, Chunbo Hao, Shuai Wang, Jixun Yao, Ziqian Ning, Meng Meng, Jian Luan, Lei Xie

机构 * School of Intelligence Science and Technology, Nanjing University, Suzhou, China(智能科学与技术学院,南京大学,苏州,中国) MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus)

专题命中 多模态生成 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20147 2025-07-25 cs.CV 57%

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

Jin Wang, Yao Lai, Aoxue Li, Shifeng Zhang, Jiacheng Sun, Ning Kang, Chengyue Wu, Zhenguo Li, Ping Luo

机构 * The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 37 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16109 2025-07-25 eess.IV cs.CV 57%

X-ray2CTPA: Leveraging Diffusion Models to Enhance Pulmonary Embolism Classification

Noa Cahan, Eyal Klang, Galit Aviram, Yiftach Barash, Eli Konen, Raja Giryes, Hayit Greenspan

机构 * Faculty of Engineering, Tel Aviv University(工程学院,特拉维夫大学) Division of Data-Driven and Digital Medicine, Department of Medicine, Icahn School of Medicine at Mount Sinai(数据驱动与数字医学部,医学部,Mount Sinai医学院) Department of Diagnostic Imaging, Sheba Medical Center, Ramat Gan, Israel(诊断影像科,Sheba 医疗中心,Ramat Gan,以色列) Department of Radiology, Icahn School of Medicine, Mount Sinai(放射科,Icahn医学院,Mount Sinai)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments preprint, project code: https://github.com/NoaCahan/X-ray2CTPA

Journal ref npj Digit. Med. 8, 439 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16623 2025-07-23 cs.CV cs.LG 57%

Automatic Fine-grained Segmentation-assisted Report Generation

Frederic Jonske, Constantin Seibold, Osman Alperen Koras, Fin Bahnsen, Marie Bauer, Amin Dada, Hamza Kalisch, Anton Schily, Jens Kleesiek

机构 * Institute for AI in Medicine, University Medicine Essen(人工智能医学研究所,埃森大学医学中心)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16240 2025-07-23 cs.CV 57%

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling

Chao Zhou, Tianyi Wei, Nenghai Yu

机构 * University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accept by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17768 2025-07-22 cs.CV 57%

R-Genie: Reasoning-Guided Generative Image Editing

Dong Zhang, Lingfeng He, Rui Yan, Fei Shen, Jinhui Tang

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Code: https://dongzhang89.github.io/RGenie.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13761 2025-07-21 cs.CL 57%

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models

Palash Nandi, Maithili Joshi, Tanmoy Chakraborty

机构 * Department of Electrical Engineering(电气工程系) Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21042 2025-07-17 cs.CR cs.AI cs.LG 57%

What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift

Jiamin Chang, Haoyang Li, Hammond Pearce, Ruoxi Sun, Bo Li, Minhui Xue

机构 * University of New South Wales \& CSIRO's Data61 Sydney Australia Macquarie University Sydney Australia University of New South Wales Sydney Australia University of Illinois at Urbana–Champaign Champaign University of New South Wales \& CSIRO's Data61 Macquarie University University of New South Wales University of Illinois at Urbana–Champaign

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Accepted to The ACM Conference on Computer and Communications Security (CCS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11522 2025-07-16 cs.CV cs.LG 57%

CATVis: Context-Aware Thought Visualization

Tariq Mehmood, Hamza Ahmad, Muhammad Haroon Shakeel, Murtaza Taj

机构 * Lahore University of Management Sciences(拉瓦尔大学管理科学学院) Forman Christian College, University(法曼基督教学院,大学) Arbisoft(阿里软件)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted at MICCAI 2025. This is the submitted version prior to peer review. The final Version of Record will appear in the MICCAI 2025 proceedings (Springer LNCS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11119 2025-07-16 cs.CV 57%

Try Harder: Hard Sample Generation and Learning for Clothes-Changing Person Re-ID

Hankun Liu, Yujian Zhao, Guanglin Niu

机构 * School of Computer Science and Engineering, Beihang University(北航计算机科学与工程学院) School of Artificial Intelligence, Beihang University(北航人工智能学院) Zhongguancun Academy(中关村学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10562 2025-07-16 cs.AI cs.CR cs.DB cs.LG 57%

SAMEP: A Secure Protocol for Persistent Context Sharing Across AI Agents

Hari Masoor

机构 * Independent Researcher(独立研究者)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments 7 pages, 4 figures, 3 implementation examples. Original work submitted as a preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10225 2025-07-16 cs.CV 57%

Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA

Zhixuan Li, Hyunse Yoon, Sanghoon Lee, Weisi Lin

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院) Department of Electrical and Electronic Engineering, Yonsei University, Korea(延世大学电子与电气工程系)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025, 17 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14171 2025-07-15 eess.IV cs.CV 57%

Guided Neural Schrödinger bridge for Brain MR image synthesis with Limited Data

Hanyeol Yang, Sunggyu Kim, Mi Kyung Kim, Yongseon Yoo, Yu-Mi Kim, Min-Ho Shin, Insung Chung, Sang Baek Koh, Hyeon Chang Kim, Jong-Min Lee

机构 * Hanyang University(翰阳大学) Hanyang University College of Medicine(翰阳大学医学院) Chonnam National University Medical School(全南国立大学医学院) Keimyung University School of Medicine(庆尚大学医学院) Yonsei Wonju College of Medicine(延世Wonju医学院) Yonsei University College of Medicine(延世大学医学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Single column, 28 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏