arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46237 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4672 篇

2203.11933 2022-10-27 cs.LG cs.CL cs.CV cs.CY 73%

A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning

Hugo Berg, Siobhan Mackenzie Hall, Yash Bhalgat, Wonsuk Yang, Hannah Rose Kirk, Aleksandar Shtedritski, Max Bain

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 17 pages, 4 figures, 7 tables. For code and trained token embeddings, see https://github.com/oxai/debias-vision-lang; Changed to use ACL layout, added joint training with comparison figure, corrected spelling and formatting errors; This paper is accepted for publication at AACL 2022, the official version of record is in the ACL Anthology

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12079 2022-10-24 cs.CL cs.CV 73%

Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies?

Mitja Nikolaus, Emmanuelle Salin, Stephane Ayache, Abdellah Fourtassi, Benoit Favre

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments To appear at EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09263 2022-10-18 cs.CV cs.CL 73%

Vision-Language Pre-training: Basics, Recent Advances, and Future Trends

Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, Jianfeng Gao

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments A survey paper/book on Vision-Language Pre-training (102 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08029 2022-09-14 cs.CV cs.AI 73%

Image Captioning for Effective Use of Language Models in Knowledge-Based Visual Question Answering

Ander Salaberria, Gorka Azkune, Oier Lopez de Lacalle, Aitor Soroa, Eneko Agirre

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Under review. 25 pages with 4 figures

Journal ref Expert Systems with Applications, Volume 212, 2023, 118669

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.01127 2022-09-07 cs.CV cs.CL 73%

VL-BEiT: Generative Vision-Language Pretraining

Hangbo Bao, Wenhui Wang, Li Dong, Furu Wei

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01056 2022-07-13 cs.CV cs.AI 73%

Counterfactually Measuring and Eliminating Social Bias in Vision-Language Pre-training Models

Yi Zhang, Junyang Wang, Jitao Sang

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.02655 2022-06-01 cs.CV cs.CL 73%

Language Models Can See: Plugging Visual Controls in Text Generation

Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, Nigel Collier

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 21 pages, 5 figures, 5 tables; (v2 adds some experimental details)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07427 2022-05-17 cs.CV cs.CL 73%

On the Complementarity of Images and Text for the Expression of Emotions in Social Media

Anna Khlyzova, Carina Silberer, Roman Klinger

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments WASSA 2022 at ACL 2022, published at https://aclanthology.org/2022.wassa-1.1/ Please cite using https://aclanthology.org/2022.wassa-1.1.bib

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.02244 2022-04-07 cs.CV cs.CL 73%

LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, Philip H. S. Torr

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12872 2022-04-06 cs.CV cs.CL 73%

Less is More: Generating Grounded Navigation Instructions from Landmarks

Su Wang, Ceslee Montgomery, Jordi Orbay, Vighnesh Birodkar, Aleksandra Faust, Izzeddin Gur, Natasha Jaques, Austin Waters, Jason Baldridge, Peter Anderson

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments CVPR 2022 Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08718 2022-03-25 cs.CV cs.CL 73%

CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, Yejin Choi

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Journal ref EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11432 2021-11-23 cs.CV cs.AI cs.LG 73%

Florence: A New Foundation Model for Computer Vision

Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, Pengchuan Zhang

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07073 2021-08-17 cs.CV cs.CL 73%

ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration

Yuhao Cui, Zhou Yu, Chunqi Wang, Zhongzhou Zhao, Ji Zhang, Meng Wang, Jun Yu

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at ACM Multimedia 2021. Code available at https://github.com/MILVLG/rosita

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.16934 2021-03-22 cs.CV cs.CL 73%

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Paper has been published in the AAAI2021 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.06165 2020-07-28 cs.CV cs.CL cs.IR cs.LG 73%

Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks

Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, Jianfeng Gao

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments ECCV 2020, Code and pre-trained models are released: https://github.com/microsoft/Oscar

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.13073 2020-07-14 cs.CV cs.CL cs.LG 73%

A Novel Attention-based Aggregation Function to Combine Vision and Language

Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06952 2026-06-19 cs.CV 版本更新 72%

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer

LaTtE-Flow: 基于层间时间步专家流的Transformer

Ying Shen, Zhiyang Xu, Jiuhai Chen, Shizhe Diao, Jiaxin Zhang, Yuguang Yao, Joy Rimchala, Ismini Lourentzou, Lifu Huang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Maryland(马里兰大学) Nvidia(英伟达) Salesforce AI Research(Salesforce AI研究) Intuit AI Research(Intuit AI研究)

专题命中 图文多模态 :multimodal(abstract,comments);multimodal foundation model(abstract);分类 cs.CV

AI总结 提出LaTtE-Flow,一种基于预训练视觉语言模型的高效统一架构,通过层间时间步专家流和条件残差注意力机制,实现图像理解与生成,生成速度提升约6倍。

Comments Unified multimodal model, Flow-matching

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20554 2026-04-06 cs.CV 72%

When Negation Is a Geometry Problem in Vision-Language Models

当否定是在视觉-语言模型中一个几何问题

Fawaz Sammani, Tzoulio Chamiti, Paul Gavrikov, Nikos Deligiannis

机构 * ETRO Department, Vrije Universiteit Brussel(布鲁塞尔自由大学ETRO系) imec Independent Researcher(独立研究员)

专题命中 图文多模态 :multimodal(abstract,comments);image-text(abstract);分类 cs.CV

AI总结 本文探讨了视觉-语言模型中否定理解的几何问题,提出基于多模态大语言模型的评估框架,并通过表示工程操控CLIP模型实现否定意识。

Comments Accepted to CVPR (Multimodal Algorithmic Reasoning Workshop) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14465 2023-10-10 cs.CV 72%

Equivariant Similarity for Vision-Language Foundation Models

Tan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, Lijuan Wang

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV;MLLM(comments)

Comments Accepted by ICCV'23 (Oral); Add evaluation on MLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05729 2023-01-02 cs.CV cs.AI cs.CL cs.LG cs.MM 72%

CLIP-TD: CLIP Targeted Distillation for Vision-Language Tasks

Zhecan Wang, Noel Codella, Yen-Chun Chen, Luowei Zhou, Jianwei Yang, Xiyang Dai, Bin Xiao, Haoxuan You, Shih-Fu Chang, Lu Yuan

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI;multimodal(comments)

Comments This paper is greatly modified and updated to be re-submitted to another conference. The new paper is under the name "Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks", https://doi.org/10.48550/arXiv.2204.10496

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00945 2026-08-12 cs.HC 版本更新 71%

Sighted by Default: Addressing Implicit Vision Assumptions in Real-Time VLM Assistance for BLV Users

默认以视觉为中心:解决面向盲人和低视力(BLV)用户的实时视觉语言模型(VLM)辅助中的隐含视觉假设问题

Yi Zhao, Siqi Wang, Qiqun Geng, Erxin Yu, Jing Li

专题命中 图文多模态 :multimodal(title)

AI总结 针对现有面向BLV用户的实时VLM辅助工具存在的以视觉为中心的默认偏见问题,提出VIA-Agent模型,其在保持与Doubao相当成功率的同时,缩短了任务时间并减少了对话轮次,提升了用户信任度。

Comments Accepted to UIST 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06679 2026-05-11 cs.LG 71%

Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding

打破幻觉:当积极与消极在多模态解码中相遇

Yubo Jiang, Yitong An, Xin Yang, Abudukelimu Wuerkaixi, Xuxin Cheng, Fengying Xie, Zhiguo Jiang, Cao Liu, Ke Zeng, Haopeng Zhang

机构 * School of Astronautics, Beihang University(北京航空航天大学航天学院) Longcat Interaction Team, Meituan(美团Longcat交互团队) Tianmushan Laboratory, Beihang University(北京航空航天大学天门山实验室)

专题命中 图文多模态 :multimodal(title)

AI总结 本文提出PND框架,通过在解码过程中引入正负对比路径,增强视觉真实性,无需重新训练即可在POPE、MME和CHAIR数据集上取得最佳性能。

Comments Accepted by CVPR 2026 (Conference on Computer Vision and Pattern Recognition). 11 pages, 5 figures. Code available at: https://github.com/JiangYubo4399/PND

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19360 2025-12-23 cs.IR 71%

Generative vector search to improve pathology foundation models across multimodal vision-language tasks

生成向量搜索以提升多模态视觉-语言任务中的病理基础模型

Markus Ekvall, Ludvig Bergenstråhle, Patrick Truong, Ben Murrell, Joakim Lundeberg

专题命中 图文多模态 :multimodal(title)

AI总结 STHLM通过生成向量搜索方法提升多模态视觉-语言任务中病理基础模型的检索性能,实现10-30%的性能提升和10倍的维度压缩

Comments 13 pages main (54 total), 2 main figures (9 total)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10963 2025-07-16 cs.HC 71%

AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multi-modal Information Between Reality and Videos

Zheng Ning, Leyang Li, Daniel Killough, JooYoung Seo, Patrick Carrington, Yapeng Tian, Yuhang Zhao, Franklin Mingzhe Li, Toby Jia-Jun Li

专题命中 图文多模态 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16496 2025-01-08 cs.LG 71%

Can Out-of-Domain data help to Learn Domain-Specific Prompts for Multimodal Misinformation Detection?

Amartya Bhattacharya, Debarshi Brahma, Suraj Nagaje Mahadev, Anmol Asati, Vikas Verma, Soma Biswas

专题命中 图文多模态 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14020 2024-07-22 q-bio.NC cs.LG 71%

NeuroBind: Towards Unified Multimodal Representations for Neural Signals

Fengyu Yang, Chao Feng, Daniel Wang, Tianye Wang, Ziyao Zeng, Zhiyang Xu, Hyoungseob Park, Pengliang Ji, Hanbin Zhao, Yuanning Li, Alex Wong

专题命中 图文多模态 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08882 2024-07-15 cs.HC 71%

Emerging Practices for Large Multimodal Model (LMM) Assistance for People with Visual Impairments: Implications for Design

Jingyi Xie, Rui Yu, He Zhang, Sooyeon Lee, Syed Masum Billah, John M. Carroll

专题命中 图文多模态 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14252 2024-02-23 cs.HC 71%

Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for Radiology

Nur Yildirim, Hannah Richardson, Maria T. Wetscherek, Junaid Bajwa, Joseph Jacob, Mark A. Pinnock, Stephen Harris, Daniel Coelho de Castro, Shruthi Bannur, Stephanie L. Hyland, Pratik Ghosh, Mercy Ranjit, Kenza Bouzid, Anton Schwaighofer, Fernando Pérez-García, Harshita Sharma, Ozan Oktay, Matthew Lungren, Javier Alvarez-Valle, Aditya Nori, Anja Thieme

专题命中 图文多模态 :multimodal(title)

Comments to appear at CHI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01963 2021-10-06 cs.CY 71%

Multimodal datasets: misogyny, pornography, and malignant stereotypes

Abeba Birhane, Vinay Uday Prabhu, Emmanuel Kahembwe

专题命中 图文多模态 :multimodal(title)

Comments 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.01860 2020-10-20 cs.LG stat.ML 71%

JECL: Joint Embedding and Cluster Learning for Image-Text Pairs

Sean T. Yang, Kuan-Hao Huang, Bill Howe

专题命中 图文多模态 :image-text(title)

Comments ICPR2020

详情

展开后加载摘要…

URL PDF HTML 收藏