arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2308.09804 2023-08-22 cs.CV cs.AI cs.CL cs.LG 67%

VL-PET: Vision-and-Language Parameter-Efficient Tuning via Granularity Control

Zi-Yuan Hu, Yanyang Li, Michael R. Lyu, Liwei Wang

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICCV 2023 (17 pages, 6 figures, 22 tables)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09351 2023-08-21 cs.CV cs.AI cs.LG cs.MM 67%

RLIPv2: Fast Scaling of Relational Language-Image Pre-training

Hangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Albanie, Yining Pan, Tao Feng, Jianwen Jiang, Dong Ni, Yingya Zhang, Deli Zhao

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to ICCV 2023. Code and models: https://github.com/JacobYuan7/RLIPv2

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12737 2023-03-23 cs.CV cs.AI cs.CL 67%

Comparing Trajectory and Vision Modalities for Verb Representation

Dylan Ebert, Chen Sun, Ellie Pavlick

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12037 2023-03-23 cs.CV cs.AI cs.LG cs.MM 67%

Causal Reasoning Meets Visual Representation Learning: A Prospective Study

Yang Liu, Yushen Wei, Hong Yan, Guanbin Li, Liang Lin

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 35 pages, 14 figures. This work has been accepted by Machine Intelligence Research. The arxiv version is kept updating by adding more novel methods, datasets and insights. The official video interpretation of this paper can be referred at https://youtu.be/2lfNaTkcTHI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.11772 2022-12-23 cs.CL cs.AI cs.SD eess.AS 67%

A Self-Adjusting Fusion Representation Learning Model for Unaligned Text-Audio Sequences

Kaicheng Yang, Ruxuan Zhang, Hua Xu, Kai Gao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07504 2022-11-15 cs.CL cs.AI cs.CV cs.IR cs.LG 67%

On Analyzing the Role of Image for Visual-enhanced Relation Extraction

Lei Li, Xiang Chen, Shuofei Qiao, Feiyu Xiong, Huajun Chen, Ningyu Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by AAAI 2023 (Student Abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14667 2022-09-30 cs.CL cs.AI cs.MM 67%

Domain-aware Self-supervised Pre-training for Label-Efficient Meme Analysis

Shivam Sharma, Mohd Khizir Siddiqui, Md. Shad Akhtar, Tanmoy Chakraborty

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted at AACL-IJCNLP 2022 main conference. 9 Pages (main content); 6 Figures; 5 Tables and an Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.00868 2022-09-28 cs.RO 67%

VIRDO: Visio-tactile Implicit Representations of Deformable Objects

Youngsun Wi, Pete Florence, Andy Zeng, Nima Fazeli

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract)

Comments This work has been accepted to ICRA 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04214 2022-07-12 cs.IR 67%

Adaptive Structural Similarity Preserving for Unsupervised Cross Modal Hashing

Liang Li, Baihua Zheng, Weiwei Sun

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

Comments Accepted to ACM Multimedia 2022 as Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10764 2022-05-24 cs.CV cs.AI cs.CL cs.CY 67%

Evidence for Hypodescent in Visual Semantic AI

Robert Wolfe, Mahzarin R. Banaji, Aylin Caliskan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To be published at ACM FAccT 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04402 2022-05-10 cs.CL cs.CV cs.MM cs.SI 67%

Detecting the Role of an Entity in Harmful Memes: Techniques and Their Limitations

Rabindra Nath Nandi, Firoj Alam, Preslav Nakov

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at CONSTRAINT 2022 (Colocated with ACL-2022), disinformation, misinformation, factuality, harmfulness, fake news, propaganda, multimodality, text, images, videos, network structure, temporality

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.02937 2022-05-09 cs.CV cs.AI cs.CL cs.LG 67%

Detection of Propaganda Techniques in Visuo-Lingual Metaphor in Memes

Sunil Gundapu, Radhika Mamidi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Paper accepted at 2nd International Conference on Machine Learning Techniques and Data Science (MLDS 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00302 2022-05-03 cs.LG 67%

SHAPE: An Unified Approach to Evaluate the Contribution and Cooperation of Individual Modalities

Pengbo Hu, Xingyu Li, Yi Zhou

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract)

Comments submitted to IJCAI2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08261 2022-04-19 cs.CV cs.AI cs.CL cs.LG q-bio.NC 67%

Visio-Linguistic Brain Encoding

Subba Reddy Oota, Jashn Arora, Vijay Rowtula, Manish Gupta, Raju S. Bapi

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 18 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06825 2022-03-25 cs.CV cs.AI cs.CL cs.LG 67%

VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks

Yi-Lin Sung, Jaemin Cho, Mohit Bansal

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2022 (15 pages; with new video-text and CLIP-ViL experiments)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.07253 2022-02-18 cs.CL cs.AI cs.LG cs.SD eess.AS 67%

Integration of Pre-trained Networks with Continuous Token Interface for End-to-End Spoken Language Understanding

Seunghyun Seo, Donghyun Kwak, Bowon Lee

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments Accepted for ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.10246 2021-09-22 cs.CL cs.AI cs.CV 67%

Does Vision-and-Language Pretraining Improve Lexical Grounding?

Tian Yun, Chen Sun, Ellie Pavlick

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Camera ready for Findings of EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01949 2021-09-07 cs.LG cs.AI cs.CL cs.CV eess.IV 67%

Improving Joint Learning of Chest X-Ray and Radiology Report by Word Region Alignment

Zhanghexuan Ji, Mohammad Abuzar Shaikh, Dana Moukheiber, Sargur Srihari, Yifan Peng, Mingchen Gao

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 10 Pages, 1 Figure, 3 Tables, Accepted in 12th Machine Learning in Medical Imaging (MLMI 2021) workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.04247 2021-08-30 cs.CV cs.AI cs.CL 67%

Reciprocal Attention Fusion for Visual Question Answering

Moshiur R Farazi, Salman H Khan

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To appear in the British Machine Vision Conference (BMVC), September 2018

Journal ref Proceedings of the British Machine Vision Conference (250) 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02605 2021-04-07 cs.CV cs.AI cs.MM 67%

An Unsupervised Sampling Approach for Image-Sentence Matching Using Document-Level Structural Information

Zejun Li, Zhongyu Wei, Zhihao Fan, Haijun Shan, Xuanjing Huang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments To be published in AAAI2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00562 2020-10-02 cs.CL cs.AI cs.CV 67%

ISAAQ -- Mastering Textbook Questions with Pre-trained Transformers and Bottom-Up and Top-Down Attention

Jose Manuel Gomez-Perez, Raul Ortega

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for publication as a long paper in EMNLP2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.09211 2020-03-23 cs.CL cs.AI eess.AS 67%

Parallel Intent and Slot Prediction using MLB Fusion

Anmol Bhasin, Bharatram Natarajan, Gaurav Mathur, Himanshu Mangla

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.04748 2020-01-14 cs.RO 67%

Expectation-Maximization for Adaptive Mixture Models in Graph Optimization

Tim Pfeifer, Peter Protzel

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract)

Comments 7 pages, 4 figures, Small fixes on Eqation. (9) and (19)

Journal ref 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 2019, pp. 3151-3157

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.07432 2016-10-25 cs.AI cs.CL cs.CV 67%

Virtual Embodiment: A Scalable Long-Term Strategy for Artificial Intelligence Research

Douwe Kiela, Luana Bulat, Anita L. Vero, Stephen Clark

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14924 2026-08-18 cs.CV cs.AI 新提交 66%

PaSTel: Anchoring Histology in Spatial Transcriptomics via Multi-Scale Hierarchical Bio-Prior Contrastive Pretraining

PaSTel:通过多尺度层级生物先验对比预训练锚定空间转录组学中的组织学

Azim Dehghani Amirabad, Junchao Zhu, Pushpak Pati, Walid Abdelmoula, Tommaso Mansi, Rui Liao

机构 * Johnson & Johnson Innovative Medicine(强生创新医药)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI;multi-modal(comments)

AI总结 PaSTel是一种整合多尺度生物先验的层级多模态预训练框架,通过三层生物先验设计解决现有空间转录组学方法的局限,在多个下游任务中性能优于现有编码器。

Comments This paper was accepted to the 3rd ICML 2026 Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13694 2024-10-18 cs.CV cs.CL 66%

Exploring the Design Space of Visual Context Representation in Video MLLMs

Yifan Du, Yuqi Huo, Kun Zhou, Zijia Zhao, Haoyu Lu, Han Huang, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, Ji-Rong Wen

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL;MLLM(comments)

Comments Long Video MLLM; work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19355 2026-08-21 cs.MM cs.CV 新提交 62%

GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

GRACE:基于适配器组合与证据感知校准的教育视觉问答的接地推理

Xinjin Li, Yudi Xia, Xi Zhao, Yiliu Xu, Yining Liu, Cheng Lu, Yujian Long, Yu Ma, Jinghan Cao, Liang Fan, Yeyun Xu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.MM

AI总结 针对教育视觉问答的问题-选项捷径问题,提出GRACE框架,利用结构化教育状态实现参数高效多模态适配,在ScienceQA上提升了多项准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18696 2026-08-20 cs.CV cs.AI 新提交 62%

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

迭代微调对复杂历史梵文手稿转录准确率的影响

Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis

机构 * Centre for Inter-disciplinary Artificial Intelligence (CAI)(跨学科人工智能中心(CAI)) FLAME University(火焰大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 针对复杂历史梵文手稿的OCR挑战,提出可在布局与外观层级迭代微调的传统OCR流水线,构建含精细标注的数据集,验证了迭代微调的性能提升并完成多模态大模型基准测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17550 2026-08-19 cs.CV cs.CL 新提交 62%

Code as Representation: A Compilable Parsing Paradigm for Academic Documents

代码作为表示:学术文档的可编译解析范式

Rihui Jin, Jun Wang, chengyuan zhu, Liang Mingyu, Yue Gao, Li Yunxuan, Kuicai Dong, Guilin Qi, Lin Ren, Yongrui Chen, Xinbang Dai, Jiaqi Li, Tongtong Wu, Gholamreza Haffari

机构 * Southeast University(东南大学) Nanyang Technological University(南洋理工大学) Nanjing University(南京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 针对学术PDF难以被机器处理的问题,提出CADP可编译解析范式,构建CADP-Bench基准并测试SOTA MLLMs,发现前沿模型仍难生成高保真可执行重构,该基准已开放供研究。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12852 2026-08-19 cs.CL cs.AI 版本更新 62%

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

AI对语言的表征中,虚假与不可能是不同的方向

Yoon Pyo Lee

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过对Gemma 3 4B IT模型的激活分析,发现其内部能区分语言的不可能性与虚假,但将偶然虚假混同于矛盾,且不可能性表征与真值、语义异常表征方向不同。

详情

展开后加载摘要…

URL PDF HTML 收藏