arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 13 篇

2511.06793 2025-11-11 cs.LG cs.AI 86%

Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models

Kunhao Li, Wenhao Li, Di Wu, Lei Yang, Jun Bai, Ju Jia, Jason Xue

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(title);分类 cs.AI

Comments Accepted at AAAI 2026 as a Conference Paper (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12658 2025-11-11 cs.DC 85%

HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving

Xianzhe Dong, Tongxuan Liu, Yuting Zeng, Liangyu Liu, Yang Liu, Siyu Wu, Yu Wu, Hailong Yang, Ke Zhang, Jing Li

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06284 2025-11-11 cs.CV cs.CL cs.MM 82%

Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective

Bing Wang, Ximing Li, Yanjun Wang, Changchun Li, Lin Yuanbo Wu, Buyu Wang, Shengsheng Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by AAAI 2026. 13 pages, 6 figures. Code: https://github.com/wangbing1416/RETSIMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10566 2025-11-11 cs.CV 79%

EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing

Umar Khalid, Kashif Munir, Hasan Iqbal, Azib Farooq, Jing Hua, Nazanin Rahnavard, Chen Chen, Victor Zhu, Zhengping Ji

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00596 2025-11-11 cs.CV 70%

Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control

Danfeng Li, Hui Zhang, Sheng Wang, Jiacheng Li, Zuxuan Wu

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) HiThink Research(HiThink研究机构)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22812 2025-11-11 cs.CL 57%

EditGRPO: Reinforcement Learning with Post-Rollout Edits for Clinically Accurate Chest X-Ray Report Generation

Kai Zhang, Christopher Malon, Lichao Sun, Martin Renqiang Min

机构 * NEC Laboratories America(NEC美国实验室) Lehigh University(莱特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18452 2025-11-11 cs.SD eess.AS 57%

DIFFA: Large Language Diffusion Models Can Listen and Understand

Jiaming Zhou, Hongjie Chen, Shiwan Zhao, Jian Kang, Jie Li, Enzhi Wang, Yujie Guo, Haoqin Sun, Hui Wang, Aobo Kong, Yong Qin, Xuelong Li

机构 * College of Computer Science, Nankai University(南开大学计算机科学学院) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究所)

专题命中 多模态生成 :multimodal(abstract);分类 eess.AS

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06422 2025-11-11 cs.CV 57%

DiffusionUavLoc: Visually Prompted Diffusion for Cross-View UAV Localization

Tao Liu, Kan Ren, Qian Chen

机构 * School of Electronic and Optical Engineering, Nanjing University of Science and Technology(电子与光学工程学院,南京理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15387 2025-11-11 cs.CV 57%

DIO: Refining Mutual Information and Causal Chain to Enhance Machine Abstract Reasoning Ability

Ruizhuo Song, Beiming Yuan

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments 15 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01322 2025-11-11 cs.CV 57%

FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors

Chenxi Li, Weijie Wang, Qiang Li, Bruno Lepri, Nicu Sebe, Weizhi Nie

机构 * Tianjin University(天津大学) University of Trento(特伦托大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

Comments Accepted by ACMMM2025, Our project webpage: https://tjulcx.github.io/FreeInsert/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05606 2025-11-11 cs.CV 57%

FreeBlend: Advancing Concept Blending with Staged Feedback-Driven Interpolation Diffusion

Yufan Zhou, Haoyu Shen, Huan Wang

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of Science and Technology of China(中国科学技术大学) Westlake University(西湖大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Webpage: https://petershen-csworld.github.io/FreeBlend

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06892 2025-11-11 cs.RO 50%

Multi-Agent AI Framework for Road Situation Detection and C-ITS Message Generation

Kailin Tong, Selim Solmaz, Kenan Mujkic, Gottfried Allmer, Bo Leng

机构 * Virtual Vehicle Research GmbH(虚拟车辆研究有限公司) ASFINAG Maut Service GmbH(ASFINAG收费服务有限公司) Tongji University(同济大学)

专题命中 多模态生成 :multimodal(abstract)

Comments submitted to TRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10684 2025-11-11 cs.LG math.OC stat.CO stat.ML 50%

MDNS: Masked Diffusion Neural Sampler via Stochastic Optimal Control

Yuchen Zhu, Wei Guo, Jaemoo Choi, Guan-Horng Liu, Yongxin Chen, Molei Tao

机构 * Georgia Institute of Technology(佐治亚理工学院) FAIR at Meta(Meta的FAIR)

专题命中 多模态生成 :multi-modal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏