arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2607.15592 2026-07-20 cs.AI 新提交 93%

MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

MGDT:具有关系自适应专家混合的MLLM引导扩散变压器用于多模态知识图谱补全

Xu Hou, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue, Wu Liu, Guanhua Ye, Lei Shi, Kangkang Lu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Communication University of China(中国传媒大学)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 研究多模态知识图谱补全问题,提出MGDT框架,先通过RASR-MoE模块选路径、抑干扰,再用MLLM对齐表示,最后KGDT去噪生成,实验证明该框架在三个基准数据集上性能优于基线。

Comments 8pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17726 2026-05-20 cs.CV cs.AI 93%

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM

Slot-MLLM: 多模态大语言模型中的面向对象视觉标记化

Donghwan Chi, Hyomin Kim, Yoonjin Oh, Yongjin Kim, Donghoon Lee, Daejin Jo, Jongmin Kim, Junyeob Baek, Sungjin Ahn, Sungwoong Kim

机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系) Kakao Corp(Kakao公司) School of Computing, KAIST(韩国科学技术院计算机学院)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种面向对象的视觉标记化方法Slot-MLLM,通过基于Slot Attention的标记器,有效编码局部视觉细节并保持高层语义,从而提升多模态大语言模型在视觉内容理解和生成中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21908 2026-07-27 cs.MM 新提交 92%

Unsupervised Multimodal Intent Discovery via MLLM-Guided Concept Generation and Semantic Propagation

通过MLLM引导的概念生成和语义传播进行无监督多模态意图发现

Yunjin Gu, Qianrui Zhou, Hua Xu

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(title,abstract);分类 cs.MM

AI总结 该研究针对无监督多模态意图发现缺乏语义监督及可解释性差的问题,提出MCSP方法,通过MLLM引导的对比推理获取语义概念,再经语义传播生成伪标签,实验证明其性能优于现有方法且能产生可解释的簇。

Comments Accepted at ACM Multimedia 2026 (MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01978 2026-07-03 cs.AI cs.CL cs.CV 新提交 92%

Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing

多模态知识编辑范围泛化用于在线递归MLLM编辑

Siyuan Li, Youyuan Zhang, Ruitong Liu, Junxi Wang, Jing Li

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Peng Cheng Laboratory(鹏城实验室) Peking University(北京大学) Fudan University(复旦大学)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 针对在线多模态知识编辑中编辑范围难以控制的问题,提出ScopeEdit方法,通过模态局部吸收分支和证据门控共享泛化分支实现范围分离的在线编辑,在保持可靠性的同时提升跨模态泛化与无关输入隔离。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15299 2026-07-20 cs.MM cs.CV cs.LG 新提交 92%

MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation

MLLM-DataEngine:闭合多模态指令微调数据生成的循环

Zhiyuan Zhao, Bin Wang, Linke Ouyang, Yiqi Lin, Pan Zhang, Xiaoyi Dong, Jiaqi Wang, Conghui He

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(title);分类 cs.CV、cs.MM

AI总结 本文提出MLLM-DataEngine闭环系统,通过自适应坏例采样模块分析模型弱点,为GPT-4提供信息以生成高质量增量数据集,能有针对性且自动地提升MLLMs能力,有望成为MLLMs数据管理通用方案。

Comments 6 pages, 4 figures, 7 tables; accepted by ICME 2026

Journal ref 2025 IEEE International Conference on Multimedia and Expo (ICME), Nantes, France, 30 June 2025 - 04 July 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16513 2026-08-18 cs.CV cs.AI 新提交 91%

MLLM-Guided Semantic Correction for Text-to-Video Generation

基于多模态大语言模型(MLLM)的文本到视频生成语义修正

Junhao Chen, Zheqi Lv, Keting Yin, Shengyu Zhang, Zhou Zhao, Feiyang Chen, Xinyu Duan, Baoxing Huai, Fei Wu

机构 * Zhejiang University(浙江大学) Huawei Cloud Computing Technology Co., Ltd.(华为云计算技术有限公司)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 该研究提出一种无需训练的MLLM引导文本到视频生成语义修正框架,通过两个关键模块在生成中修正语义偏差,提升了生成内容的语义对齐度等性能,经多基准实验验证有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16484 2026-08-18 cs.CV 新提交 91%

Remote-Sensing City Layout Extraction with MLLM

基于多模态大语言模型(MLLM)的遥感城市布局提取

Zigan Zhou, Kai Li, Yupeng Deng

机构 * City University of Hong Kong(香港城市大学) University of Chinese Academy of Sciences(中国科学院大学) Aerospace Information Research Institute(空天信息创新研究院)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV

AI总结 本研究提出Code-as-City方法,利用MLLM将遥感顶视图图像转化为含平面与3D输出的可编辑城市布局,在CityLayout-100数据集上取得41.1%、48.3%的交并比,验证了视觉观测转城市代码的可行性。

Comments 4 pages, 2 figures, 4 tables. Accepted to IEEE APGARSS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.04128 2026-05-21 cs.GR cs.AI cs.CL cs.CV cs.LG 90%

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

JoyAI-Image: 激活统一多模态理解和生成中的空间智能

Lin Song, Wenbo Li, Guoqing Ma, Wei Tang, Bo Wang, Yuan Zhang, Yijun Yang, Yicheng Xiao, Jianhui Liu, Yanbing Zhang, Guohui Zhang, Wenhu Zhang, Hang Xu, Nan Jiang, Xin Han, Haoze Sun, Maoquan Zhang, Haoyang Huang, Nan Duan

机构 * Joy Future Academy, JD(京东探索研究院)

专题命中 多模态生成 :MLLM(summary_cn,abstract);multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出JoyAI-Image,一种统一的多模态基础模型,用于视觉理解、文本到图像生成和指令引导的图像编辑。该模型结合了空间增强的多模态大语言模型(MLLM)和多模态扩散Transformer(MMDiT),通过共享的多模态接口实现感知与生成的交互。构建可扩展的训练配方,结合统一指令微调、长文本渲染监督、空间 grounded 数据和通用及空间编辑信号,使模型具备广泛的多模态能力,同时增强几何感知推理和可控视觉合成。实验表明,JoyAI-Image在理解、生成、长文本渲染和编辑基准上达到最先进的性能。更重要的是,增强的理解、可控的空间编辑和新视角辅助推理之间的双向循环使模型超越一般视觉能力,向更强的空间智能发展。

Comments Code: https://github.com/jd-opensource/JoyAI-Image

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16408 2026-06-16 cs.LG 新提交 90%

MUNI: Multimodal Unified Latent Diffusion for Coherent Any-to-Any Generation

MUNI:面向连贯任意到任意生成的多模态统一潜在扩散

Kyeongmin Yeo, Yunhong Min, Minhyuk Sung

机构 * KAIST(韩国科学技术院)

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(title,abstract);cross-modal(abstract);image-text(abstract)

AI总结 提出MUNI框架,通过端到端多模态潜在扩散和路由训练目标,实现任意到任意生成,在条件生成上匹配或超越基线,并在无条件连贯性上取得最大优势。

Comments Project page: https://muni-proj.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14043 2026-08-17 cs.CV 新提交 90%

Beyond Text Conditioning: A Systematic Study of MLLM-DiT Fusion for Video Generation

超越文本条件:面向视频生成的MLLM-DiT融合系统研究

Yanbo Ding, Yijia Fan, Caihua Shan, Yifan Yang, Yifei Shen, Weijie Wang, Xirui Hu, Dongsheng Li, Lili Qiu, Yuqing Yang, Yali Wang

机构 * Chinese Academy of Sciences(中国科学院) Microsoft Research(微软研究院) Sun Yat-sen University(中山大学) Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Xi’an Jiaotong University(西安交通大学)

专题命中 多模态生成 :MLLM(title,title_cn);分类 cs.CV

AI总结 该研究针对视频生成中MLLM与DiT融合问题,提出BiVidGen框架,通过MLLM生成语义视觉标记辅助DiT渲染,提升了视频的语义对齐度与时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01955 2026-05-04 cs.CV cs.AI cs.LG 90%

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

GPT-4o对视觉的理解有多好?在标准计算机视觉任务上评估多模态基础模型

Rahul Ramachandran, Ali Garjani, Roman Bachmann, Andrei Atanov, Oğuzhan Fatih Kar, Amir Zamir

机构 * Swiss Federal Institute of Technology(瑞士联邦理工学院)

专题命中 多模态生成 :multimodal(title,abstract);multimodal foundation model(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文评估了GPT-4o等多模态基础模型在标准计算机视觉任务上的表现,发现其在语义任务上优于几何任务,GPT-4o在非推理模型中表现最佳,推理模型在几何任务中有所提升。

Comments ICLR 2026. Project page at https://fm-vision-evals.epfl.ch/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12711 2025-05-21 cs.CV cs.AI 90%

Any-to-Any Learning in Computational Pathology via Triplet Multimodal Pretraining

Qichen Sun, Zhengrui Guo, Rui Peng, Hao Chen, Jinzhuo Wang

机构 * Peking University(北京大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05519 2024-06-26 cs.AI cs.CL cs.LG 90%

NExT-GPT: Any-to-Any Multimodal LLM

Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, Tat-Seng Chua

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments ICML 2024 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18710 2026-06-18 cs.CR 新提交 89%

Image Prompt Reconstruction Attacks on Distributed MLLM Inference Frameworks

分布式多模态大模型推理框架上的图像提示重建攻击

Xinjian Luo, Hongyan Chang, Jianxin Wei, Yuncheng Wu, Xiaofeng Gao, Meikang Qiu, Ting Yu, Xue Liu

专题命中 多模态生成 :MLLM(title,summary_cn);multimodal(abstract,abstract_cn)

AI总结 研究分布式MLLM推理中中间嵌入泄露图像提示的风险,提出两种被动黑盒攻击方法MPAA和IEDA,实现像素级和语义级图像重建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07986 2025-07-24 cs.CV 89%

Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

Zhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen, Wangmeng Zuo, Ziwei Liu, Kwan-Yee K. Wong

机构 * The University of Hong Kong(香港大学) Nanjing University(南京大学) University of Chinese Academy of Sciences(中国科学院大学) Nanyang Technological University(南洋理工大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2025; Project Page: https://vchitect.github.io/TACA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18304 2024-05-29 cs.CV 89%

Multi-modal Generation via Cross-Modal In-Context Learning

Amandeep Kumar, Muzammal Naseer, Sanath Narayan, Rao Muhammad Anwer, Salman Khan, Hisham Cholakkal

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12118 2026-04-29 cs.LG cs.DC 89%

Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models

Cornserve:一种用于任意到任意多模态模型的分布式服务系统

Jae-Won Chung, Jeff J. Ma, Jisang Ahn, Yizhuo Liang, Akshay Jajoo, Myungjin Lee, Mosharaf Chowdhury

机构 * University of Michigan(密歇根大学) University of Southern California(南加州大学) Cisco Research(思科研究)

专题命中 多模态生成 :any-to-any(title,abstract);multimodal(title,abstract)

AI总结 本文提出Cornserve,一种支持任意到任意多模态模型的分布式服务系统,通过灵活的任务抽象和组件解耦实现高效部署,提升了吞吐量和延迟性能。

Comments CAIS 2026 Demo track | Open source at https://github.com/cornserve-ai/cornserve | Demo video at https://www.youtube.com/watch?v=nb8R-vztLRg

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13719 2026-03-31 cs.CV cs.AI cs.LG cs.MM cs.RO 89%

Scaling Spatial Intelligence with Multimodal Foundation Models

通过多模态基础模型提升空间智能

Zhongang Cai, Ruisi Wang, Chenyang Gu, Fanyi Pu, Junxiang Xu, Yubo Wang, Wanqi Yin, Zhitao Yang, Chen Wei, Qingping Sun, Tongxi Zhou, Jiaqi Li, Hui En Pang, Oscar Qian, Yukun Wei, Zhiqian Lin, Xuanke Shi, Kewang Deng, Xiaoyang Han, Zukai Chen, Xiangyu Fan, Hanming Deng, Lewei Lu, Liang Pan, Bo Li, Ziwei Liu, Quan Wang, Dahua Lin, Lei Yang

机构 * SenseTime Research(商汤科技研究院) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文通过构建SenseNova-SI家族,利用八百万多样数据样本提升空间智能,在多个基准测试中取得优异成绩,并分析数据扩展的影响及潜在应用。

Comments Codebase: https://github.com/OpenSenseNova/SenseNova-SI ; Models: https://huggingface.co/collections/sensenova/sensenova-si . This report is based on the v1.1 version of SenseNova-SI. Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01169 2025-03-24 cs.MM cs.CV cs.SD eess.AS 89%

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Zichun Liao, Yusuke Kato, Kazuki Kozuka, Aditya Grover

专题命中 多模态生成 :multi-modal(title,abstract);any-to-any(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments 19 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06135 2024-07-09 cs.CL cs.AI cs.CV 89%

ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Ethan Chern, Jiadi Su, Yan Ma, Pengfei Liu

专题命中 多模态生成 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15846 2026-07-30 cs.CV 版本更新 89%

A Closer Look at Dynamic Scene Graph Generation In the Era of Multimodal Large Language Models

多模态大语言模型时代的动态场景图生成研究

Xuanming Cui, Jaiminkumar Ashokbhai Bhoi, Chionh Wei Peng, Adriel Kuek, Ser Nam Lim

专题命中 多模态生成 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.CV

AI总结 本研究针对多模态大语言模型时代动态场景图生成的局限,从任务设置优化评估指标、基于MLLM改进模型设计,在多个数据集上实现了最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05981 2026-07-16 cs.CV cs.LG 版本更新 89%

Inverting the Streaming-Diffusion Bottleneck: Video-Rate MLLM-Conditioned Edit Diffusion on a Consumer GPU

基于视觉感知的多模态大语言模型条件编辑扩散的视频率流式风格化:蒸馏UNet + MLLM文本编码器上的非对称批处理推理

Yoshiyuki Ootani

机构 * Independent researcher(独立研究员)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV

AI总结 针对蒸馏扩散模型中文本编码器成为瓶颈的问题,提出一种结合非对称CUDA流水线、编译友好的ControlNet-LLLite重构和周期性条件刷新调度的流式管线,在消费级GPU上实现视频率实时风格化编辑。

Comments 14 pages, 4 figures, 13 tables. Code, evaluation harness, and the released Temporal LLLite adapter weights are at https://github.com/otanl/dreamlite-stream (also mirrored to Hugging Face and Zenodo)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29689 2026-06-30 cs.CL 89%

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

MLLM 能否像人类一样进行批评?评估多模态大语言模型中的开放式审美推理

Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani

机构 * UCSB(加州大学圣塔芭芭拉分校) University of Washington(华盛顿大学) UCLA(加州大学洛杉矶分校)

专题命中 多模态生成 :MLLM(title_cn,abstract);multimodal(title,abstract);分类 cs.CL

AI总结 本文评估多模态大语言模型在开放式审美批评中的表现,发现基于参考的相似度指标存在误导,模型在选择性、特异性和多样性方面与人类批评存在系统性差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27374 2026-05-28 cs.CL 89%

ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment

ICG: 通过基于MLLM的提示和个性化偏好对齐改进封面图像生成

Zhipeng Bian, Jieming Zhu, Qijiong Liu, Wang Lin, Guohao Cai, Zhaocheng Du, Jiacheng Sun, Zhou Zhao, Zhenhua Dong

机构 * Huazhong University of Science and Technology(华中科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Hong Kong Polytechnic University(香港理工大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CL

AI总结 提出ICG框架,利用多模态大语言模型和扩散模型,通过元标记提取语义特征、用户嵌入个性化对齐及多奖励学习策略,实现高质量、个性化封面图像生成。

Comments Published in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 12268-12278, EMNLP 2025. Official version: https://doi.org/10.18653/v1/2025.emnlp-main.617

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (Main Track) EMNLP 2025 12268-12278

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13721 2025-10-17 cs.CL cs.AI cs.CV cs.MM 89%

NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching

Run Luo, Xiaobo Xia, Lu Wang, Longze Chen, Renke Shan, Jing Luo, Min Yang, Tat-Seng Chua

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) University of Chinese Academy of Sciences(中国科学院大学) NExT++ Research Center(NExT++研究中心) National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :any-to-any(title,abstract);multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18775 2023-12-01 cs.CV cs.AI cs.CL cs.LG cs.SD eess.AS 89%

CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

Zineng Tang, Ziyi Yang, Mahmoud Khademi, Yang Liu, Chenguang Zhu, Mohit Bansal

专题命中 多模态生成 :any-to-any(title,abstract);multimodal(abstract);MLLM(abstract);multimodal foundation model(abstract)

Comments Project Page: https://codi-2.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26111 2026-05-26 cs.CV cs.AI cs.GR cs.LG cs.MM 89%

Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation

从多模态大语言模型中榨取能力用于主题驱动生成

Shuhong Zheng, Aashish Kumar Misraa, Yu-Teng Li, Yu-Jhe Li, Igor Gilitschenski

机构 * University of Toronto & Vector Institute(多伦多大学及向量研究所) Adobe(Adobe公司) Google(谷歌公司)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract,abstract_cn);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 提出一种结合多模态大语言模型和VAE身份条件的方法,通过双层级聚合模块和多阶段去噪策略,在主题驱动图像生成中实现多模态理解与身份保持的平衡,优于现有方法。

Comments 33 pages, 18 figures, Project Page: https://zsh2000.github.io/squeeze-mllm-subject-gen/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10020 2023-09-20 cs.CV cs.CL 88%

Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, Jianfeng Gao

专题命中 多模态生成 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.CL

Comments 119 pages, PDF file size 58MB; Tutorial website: https://vlp-tutorial.github.io/2023/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04015 2026-08-04 eess.SP 版本更新 88%

GenED-SC: Generative Editing Semantic Communication with Integrated Multi-Modal LLMs

GenED-SC:集成多模态大模型的生成式编辑语义通信

Shuoyao Wang, Suzhi Bi, Mingze Gong, Zhanpeng Wang, Li Ping Qian, Qiang Ye

专题命中 多模态生成 :MLLM(summary_cn,abstract);multi-modal(title);multimodal(abstract)

AI总结 提出一种两阶段语义图像传输框架,结合JSCC判别传输与MLLM生成编辑,在低信噪比下提升语义保真度和感知质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25901 2026-07-29 cs.IR 新提交 88%

RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation

RecoReward:用于推荐的推荐器引导多模态描述生成

Guohong Mu, Yueyang Liu, Jiangxia Cao, Changxin Lao, Zijie Zhuang, Yuhui Zhang, Jiaqi Feng, Ruochen Yang, Shuang Yang, Zhaojie Liu, Qibin Hou

专题命中 多模态生成 :MLLM(summary_cn,abstract);multimodal(title,abstract)

AI总结 研究针对多模态推荐中传统方法不足,提出RecoReward,训练时用行为衍生奖励,保留仅内容推理。在直播推荐中利用用户历史等估计亲和力,通过推荐器亲和力分数提供反馈,实验表明该方法能提升MLLM性能,利于下游推荐且保留内容服务。

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏