arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3484 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3484 篇

2310.07931 2023-10-13 cs.LG cs.AI cs.CL cs.CV 67%

D2 Pruning: Message Passing for Balancing Diversity and Difficulty in Data Pruning

Adyasha Maharana, Prateek Yadav, Mohit Bansal

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 17 pages (Our code is available at https://github.com/adymaharana/d2pruning)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13923 2023-08-08 cs.CV cs.CL cs.MM 67%

Retrieval-based Knowledge Augmented Vision Language Pre-training

Jiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou, Yuedong Yang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments arXiv admin note: text overlap with arXiv:2210.09338 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16395 2023-08-01 cs.CV cs.AI cs.CL cs.LG 67%

Bridging the Gap: Exploring the Capabilities of Bridge-Architectures for Complex Visual Reasoning Tasks

Kousik Rajesh, Mrigank Raman, Mohammed Asad Karim, Pranit Chawla

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05665 2023-06-01 cs.CV cs.AI cs.LG cs.MM 67%

ImageBind: One Embedding Space To Bind Them All

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments CVPR 2023 (Highlighted Paper). Website: https://imagebind.metademolab.com/ Code/Models: https://github.com/facebookresearch/ImageBind

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12646 2023-02-21 cs.IR 67%

MAKE: Vision-Language Pre-training based Product Retrieval in Taobao Search

Xiaoyang Zheng, Zilong Wang, Ke Xu, Sen Li, Tao Zhuang, Qingwen Liu, Xiaoyi Zeng

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)

Comments 5 pages, accepted to The Industry Track of the Web Conference 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04269 2023-02-09 cs.LG cs.AI cs.CL cs.CV 67%

Diagnosing and Rectifying Vision Models using Language

Yuhui Zhang, Jeff Z. HaoChen, Shih-Cheng Huang, Kuan-Chieh Wang, James Zou, Serena Yeung

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07094 2023-01-18 cs.CV cs.AI cs.CL cs.LG 67%

Learning Customized Visual Models with Retrieval-Augmented Knowledge

Haotian Liu, Kilho Son, Jianwei Yang, Ce Liu, Jianfeng Gao, Yong Jae Lee, Chunyuan Li

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02495 2022-09-16 cs.CV cs.AI cs.CL 67%

Sign Language Video Retrieval with Free-Form Textual Queries

Amanda Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül Varol

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01058 2022-07-05 cs.AI cs.CV cs.HC cs.MM 67%

Chat-to-Design: AI Assisted Personalized Fashion Design

Weiming Zhuang, Chongjie Ye, Ying Xu, Pengzhi Mao, Shuai Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10787 2022-02-23 cs.CL cs.AI cs.CV cs.LG 67%

VU-BERT: A Unified framework for Visual Dialog

Tong Ye, Shijing Si, Jianzong Wang, Rui Wang, Ning Cheng, Jing Xiao

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 5 pages, 2 figures, accepted by 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07074 2021-12-06 cs.LG cs.AI cs.CL cs.CV 67%

Memotion Analysis through the Lens of Joint Embedding

Nethra Gunti, Sathyanarayanan Ramamoorthy, Parth Patwa, Amitava Das

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted as Student Abstract at AAAI-22

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09150 2021-10-22 cs.CL cs.AI cs.CV 67%

VisualSem: A High-quality Knowledge Graph for Vision and Language

Houda Alberts, Teresa Huang, Yash Deshpande, Yibo Liu, Kyunghyun Cho, Clara Vania, Iacer Calixto

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for publication at the 1st Multilingual Representation Learning workshop (MRL 2021) co-located with EMNLP 2021. 15 pages, 8 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.09109 2021-02-19 cs.CV cs.AI cs.MM 67%

Understanding and Creating Art with AI: Review and Outlook

Eva Cetinic, James She

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 17 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.15086 2021-01-01 cs.CL cs.AI cs.CV 67%

Accurate Word Representations with Universal Visual Guidance

Zhuosheng Zhang, Haojie Yu, Hai Zhao, Rui Wang, Masao Utiyama

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.03687 2020-05-26 cs.LG stat.ML 67%

COBRA: Contrastive Bi-Modal Representation Algorithm

Vishaal Udandarao, Abhishek Maiti, Deepak Srivatsav, Suryatej Reddy Vyalla, Yifang Yin, Rajiv Ratn Shah

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract)

Comments 13 Pages, 6 Figures and 10 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.11449 2020-04-27 cs.MM cs.CL cs.CV cs.LG 67%

Upgrading the Newsroom: An Automated Image Selection System for News Articles

Fangyu Liu, Rémi Lebret, Didier Orel, Philippe Sordet, Karl Aberer

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted to ACM Transactions on Multimedia Computing Communications and Applications (ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.00850 2019-11-05 cs.AI cs.CL cs.CV 67%

Scene Graph based Image Retrieval -- A case study on the CLEVR Dataset

Sahana Ramnath, Amrita Saha, Soumen Chakrabarti, Mitesh M. Khapra

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 3 pages including references, Accepted at the ICCV 2019 Workshop - 'Linguistics Meets Image and Video Retrieval' (received Best Paper Award)

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.08454 2018-12-27 cs.CL cs.AI cs.CV cs.LG 67%

Attention Based Natural Language Grounding by Navigating Virtual Environment

Akilesh B, Abhishek Sinha, Mausoom Sarkar, Balaji Krishnamurthy

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at WACV 2019. Also at NeurIPS 2017 workshop on Visually-Grounded Interaction and Language (ViGIL)

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.01606 2018-07-02 cs.MM cs.AI cs.CV cs.LG 67%

Multimedia Semantic Integrity Assessment Using Joint Embedding Of Images And Text

Ayush Jaiswal, Ekraam Sabir, Wael AbdAlmageed, Premkumar Natarajan

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments *Ayush Jaiswal and Ekraam Sabir contributed equally to the work in this paper

Journal ref In Proceedings of the 2017 ACM on Multimedia Conference, pp. 1465-1471. ACM, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.02717 2016-08-10 cs.CV cs.AI cs.CL cs.LG 67%

Mean Box Pooling: A Rich Image Representation and Output Embedding for the Visual Madlibs Task

Ashkan Mokarian, Mateusz Malinowski, Mario Fritz

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to BMVC'16

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.05659 2015-11-19 cs.IR 67%

Learning Discriminative Representations for Semantic Cross Media Retrieval

Aiwen Jiang, Hanxi Li, Yi Li, Mingwen Wang

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12295 2026-06-11 cs.CV cs.CL cs.IR 新提交 66%

Findings of the MAGMaR 2026 Shared Task

MAGMaR 2026 共享任务结果

Alexander Martin, Dengjia Zhang, Joel Brogan, Francis Ferraro, Jeremy Gwinnup, Reno Kriz, Teng Long, Kenton Murray, Andrew Yates, Xiang Xiang

机构 * Johns Hopkins University(约翰霍普金斯大学) OpenAI University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) Air Force Research Laboratory(空军研究实验室) Human Language Technology Center of Excellence, Johns Hopkins University(约翰霍普金斯大学人类语言技术卓越中心) University of Amsterdam(阿姆斯特丹大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 跨模态检索 :multimodal(abstract,comments);分类 cs.CV、cs.CL

AI总结 本文介绍MAGMaR 2026共享任务的结果,包括视频检索和基于检索视频的生成任务,所有提交系统均超越去年基线。

Comments Findings of the 2nd workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR); Resources at this url: https://github.com/rekriz11/MAGMAR_2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15735 2026-04-20 cs.CV cs.AI 66%

Sketch and Text Synergy: Fusing Structural Contours and Descriptive Attributes for Fine-Grained Image Retrieval

草图与文本协同:融合结构轮廓和描述属性用于细粒度图像检索

Siyuan Wang, Hanchen Gao, Guangming Zhu, Jiang Lu, Yiyue Ma, Tianci Wu, Jincai Huang, Liang Zhang

机构 * Xidian University First Aircraft Design Institute(西安电子科技大学第一飞机设计院)

专题命中 跨模态检索 :cross-modal(abstract,comments);分类 cs.CV、cs.AI

AI总结 本文提出STBIR框架,通过融合草图的结构轮廓与文本的色彩纹理信息,提升细粒度图像检索性能,采用课程学习、特征空间优化和多阶段跨模态对齐机制,实验验证其优于现有方法。

Comments Image Retrieval, Hand-drawn Sketch, Multi-stage Cross-modal Feature Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20670 2025-06-26 cs.CV cs.CL 66%

MMSearch-R1: Incentivizing LMMs to Search

Jinming Wu, Zihao Deng, Wei Li, Yiding Liu, Bo You, Bo Li, Zejun Ma, Ziwei Liu

机构 * ByteDance(字节跳动) S-Lab, NTU(NTU的S实验室)

专题命中 跨模态检索 :multimodal(abstract,comments);分类 cs.CV、cs.CL

Comments Code: https://github.com/EvolvingLMMs-Lab/multimodal-search-r1

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17706 2024-05-29 cs.AI cs.CV cs.IR 66%

Video Enriched Retrieval Augmented Generation Using Aligned Video Captions

Kevin Dela Rosa

专题命中 跨模态检索 :multimodal(abstract,comments);分类 cs.CV、cs.AI

Comments SIGIR 2024 Workshop on Multimodal Representation and Retrieval (MRR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21431 2026-08-25 cs.CV cs.MM 新提交 62%

Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

基于结构化上下文推理提升基于知识的视觉问答性能

Qiyou Liu, Yong Zhang, Jianjie Luo, Zhenguo Yang, Yi Yu

机构 * School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Graduate School of Advanced Science and Engineering, Hiroshima University(广岛大学先进科学与工程研究生院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

AI总结 本文提出SCoRe框架,通过上下文获取、选择、压缩三阶段处理多模态知识,在OK-VQA和A-OKVQA基准上性能优于现有最优方法,提升了基于知识的视觉问答效果。

Comments Accepted by ICME 2026. Source code is available at https://github.com/WISLab-GDUT/SCoRe

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12746 2026-08-20 cs.CV cs.CL 版本更新 62%

Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors

双流跨锚校正:接地长文本描述与对象级锚点的域限制

Lingkai Bu, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao, Jinyi Liang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 针对多模态大语言模型的长文本描述对象幻觉问题,本文提出双流跨锚校正方法,通过耦合感知流与认知流提升精度,在长文本场景下实现最优性能,且存在域条件性限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12779 2026-08-18 cs.CV cs.CL 62%

Open Vocabulary Panoptic Segmentation With Retrieval Augmentation

开放词汇全景分割与检索增强

Nafis Sadeq, Qingfeng Liu, Mostafa El-Khamy

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL

AI总结 本文提出RetCLIP方法,通过检索增强提升开放词汇全景分割的性能,实现对未见类别的有效分割。

Journal ref IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14252 2026-08-17 cs.AI cs.CL 新提交 62%

Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models

无校正控制的基础:大语言模型的真值追踪剖面

Brett Reynolds

机构 * Humber Polytechnic(亨伯理工学院) University of Toronto(多伦多大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文研究大语言模型中无校正控制的基础问题,提出路径剖面概念以分析真值追踪,指出纯文本模型继承的模式可提供衍生可应答性,不同方法对任务的真值追踪改进可能与表面改进不一致。

Comments 24 pages, 1 figure, 1 table. A six-page methodological supplement, reproducible R script, and constructed data are included as ancillary files

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13939 2026-08-17 cs.CV cs.AI 新提交 62%

CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

CMCNet:将超声图像嵌入与文本TI-RADS表示对齐以用于细粒度甲状腺分类

Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究构建含600个甲状腺结节的STN数据集,提出CMCNet模型,通过中心间隔对比损失对齐图像与文本嵌入,实现细粒度甲状腺分类,性能优于多类基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏