arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2410.02761 2025-04-15 cs.CV cs.AI 81%

FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models

Zhipei Xu, Xuanyu Zhang, Runyi Li, Zecheng Tang, Qing Huang, Jian Zhang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11079 2025-04-11 cs.CV cs.AI 81%

Phantom: Subject-consistent video generation via cross-modal alignment

Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Jiawei Liu, Gen Li, Siyu Zhou, Qian He, Xinglong Wu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05314 2025-04-09 cs.IR cs.AI cs.CL 81%

Multimodal Quantitative Language for Generative Recommendation

Jianyang Zhai, Zi-Feng Mai, Chang-Dong Wang, Feidiao Yang, Xiawu Zheng, Hui Li, Yonghong Tian

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03295 2025-04-07 cs.CL cs.AI 81%

Stance-Driven Multimodal Controlled Statement Generation: New Dataset and Task

Bingqian Wang, Quan Fang, Jiachen Sun, Xiaoxiao Ma

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02312 2025-04-04 cs.CV cs.AI 81%

OmniCam: Unified Multimodal Video Generation via Camera Control

Xiaoda Yang, Jiayang Xu, Kaixuan Luan, Xinyu Zhan, Hongshun Qiu, Shijun Shi, Hao Li, Shuai Yang, Li Zhang, Checheng Yu, Cewu Lu, Lixin Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23125 2025-04-01 cs.CV cs.AI 81%

Evaluating Compositional Scene Understanding in Multimodal Generative Models

Shuhao Fu, Andrew Jun Lee, Anna Wang, Ida Momennejad, Trevor Bihl, Hongjing Lu, Taylor W. Webb

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21193 2025-03-28 cs.CL cs.CV 81%

UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning

Hongxuan Tang, Hao Liu, Xinyan Xiao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20853 2025-03-28 cs.CV cs.AI cs.LG cs.RO 81%

Unified Multimodal Discrete Diffusion

Alexander Swerdlow, Mihir Prabhudesai, Siddharth Gandhi, Deepak Pathak, Katerina Fragkiadaki

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Project Website: https://unidisc.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20118 2025-03-27 cs.GR cs.AI cs.CV 81%

Zero-Shot Human-Object Interaction Synthesis with Multimodal Priors

Yuke Lou, Yiming Wang, Zhen Wu, Rui Zhao, Wenjia Wang, Mingyi Shi, Taku Komura

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16537 2025-03-24 cs.CL cs.CV 81%

Do Multimodal Large Language Models Understand Welding?

Grigorii Khvatskii, Yong Suk Lee, Corey Angst, Maria Gibbs, Robert Landers, Nitesh V. Chawla

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12927 2025-03-20 cs.CV cs.AI 81%

MMLNB: Multi-Modal Learning for Neuroblastoma Subtyping Classification Assisted with Textual Description Generation

Huangwei Chen, Yifei Chen, Zhenyu Yan, Mingyang Ding, Chenlei Li, Zhu Zhu, Feiwei Qin

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10324 2025-03-14 cs.CV cs.MM 81%

IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification

Yuhao Wang, Yongfeng Lv, Pingping Zhang, Huchuan Lu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments This work is accepted by CVPR2025. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08133 2025-03-12 cs.CV cs.AI 81%

MGHanD: Multi-modal Guidance for authentic Hand Diffusion

Taehyeon Eum, Jieun Choi, Tae-Kyun Kim

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15191 2025-03-12 cs.CV cs.LG cs.SD eess.AS 81%

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

Moayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov, Alper Canberk, Kwot Sin Lee, Vicente Ordonez, Sergey Tulyakov

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、eess.AS

Comments Project Page: snap-research.github.io/AVLink/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18459 2025-03-04 cs.CV cs.MM 81%

FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation

Yuki Imajuku, Yoko Yamakata, Kiyoharu Aizawa

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 15 pages, 5 figures. We found errors in the calculation of evaluation metrics, which were corrected in this version with $\color{blue}{\text{modifications highlighted in blue}}$. Please also see the Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00171 2025-03-04 cs.CV cs.AI 81%

PaliGemma-CXR: A Multi-task Multimodal Model for TB Chest X-ray Interpretation

Denis Musinguzi, Andrew Katumba, Sudi Murindanyi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20172 2025-02-28 cs.CV cs.CL 81%

Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think

Liang Chen, Shuai Bai, Wenhao Chai, Weichu Xie, Haozhe Zhao, Leon Vinci, Junyang Lin, Baobao Chang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 13 pages, 9 figures, codebase in https://github.com/chenllliang/DreamEngine

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11829 2025-02-18 cs.CL cs.AI cs.SE 81%

Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities

Hanbin Wang, Xiaoxuan Zhou, Zhipeng Xu, Keyuan Cheng, Yuxin Zuo, Kai Tian, Jingwei Song, Junting Lu, Wenhui Hu, Xueyang Liu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10536 2025-02-18 cs.CV cs.AI cs.LG 81%

PolyPath: Adapting a Large Multimodal Model for Multi-slide Pathology Report Generation

Faruk Ahmed, Lin Yang, Tiam Jaroensri, Andrew Sellergren, Yossi Matias, Avinatan Hassidim, Greg S. Corrado, Dale R. Webster, Shravya Shetty, Shruthi Prabhakara, Yun Liu, Daniel Golden, Ellery Wulczyn, David F. Steiner

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 main pages, 21 pages in total

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22108 2025-02-18 cs.CL cs.AI 81%

Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench

Zheyuan Liu, Guangyao Dou, Mengzhao Jia, Zhaoxuan Tan, Qingkai Zeng, Yongle Yuan, Meng Jiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments NAACL Main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09843 2025-02-17 cs.AI cs.HC cs.MM 81%

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System

Karan Taneja, Ashok K. Goel

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments 5 pages, 3 figures, AAAI-MAKE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03163 2025-02-11 cs.CL cs.CV cs.CY 81%

Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

Chenglei Si, Yanzhe Zhang, Ryan Li, Zhengyuan Yang, Ruibo Liu, Diyi Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments NAACL 2025; The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03429 2025-02-06 cs.CL cs.AI 81%

On Fairness of Unified Multimodal Large Language Model for Image Generation

Ming Liu, Hao Chen, Jindong Wang, Liwen Wang, Bhiksha Raj Ramakrishnan, Wensheng Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14170 2025-02-05 cs.IR cs.AI cs.MM 81%

Personalized Image Generation with Large Multimodal Models

Yiyan Xu, Wenjie Wang, Yang Zhang, Biao Tang, Peng Yan, Fuli Feng, Xiangnan He

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted for publication in WWW'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16698 2025-01-29 cs.CL cs.CV cs.RO 81%

3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow

Yueen Ma, Yuzheng Zhuang, Jianye Hao, Irwin King

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Preprint. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15393 2025-01-28 cs.AI cs.CL 81%

Diffusion-based Hierarchical Negative Sampling for Multimodal Knowledge Graph Completion

Guanglin Niu, Xiaowei Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments The version of a full paper accepted to DASFAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18216 2025-01-22 cs.CV cs.CL 81%

ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation

Mengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu, Zhongming Jiang, Xuehui Wang, Qinbin Li, Guangneng Hu, Shengchao Qin, Chi-Wing Fu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by the AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08187 2025-01-16 cs.CL cs.AI cs.CE cs.HC cs.LG q-bio.CB 81%

A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following

Yin Fang, Xinle Deng, Kangwei Liu, Ningyu Zhang, Jingyang Qian, Penghui Yang, Xiaohui Fan, Huajun Chen

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 37 pages; 13 figures; Code: https://github.com/zjunlp/Instructcell, Models: https://huggingface.co/zjunlp/Instructcell-chat, https://huggingface.co/zjunlp/InstructCell-instruct

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02523 2025-01-07 cs.CV cs.AI 81%

Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation

Dawei Dai, Mingming Jia, Yinxiu Zhou, Hang Xing, Chenghang Li

专题命中 多模态生成 :multimodal(title);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06550 2025-01-07 cs.CV cs.AI 81%

Multimodal Urban Areas of Interest Generation via Remote Sensing Imagery and Geographical Prior

Chuanji Shi, Yingying Zhang, Jiaotuan Wang, Xin Guo, Qiqi Zhu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 9 figures

Journal ref International Journal of Applied Earth Observation and Geoinformation, 136(2025)

详情

展开后加载摘要…

URL PDF HTML 收藏