arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3475 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3475 篇

2506.17302 2025-06-24 cs.CV cs.LG 79%

Fine-Scale Soil Mapping in Alaska with Multimodal Machine Learning

Yijun Lin, Theresa Chen, Colby Brungard, Grunwald Sabine, Sue Ives, Matt Macander, Timm Nawrocki, Yao-Yi Chiang, Nic Jelinski

机构 * University of Minnesota(明尼苏达大学) New Mexico State University(新墨西哥州立大学) University of Florida(佛罗里达大学) ABR, Inc.(ABR公司) University of Alaska-Anchorage(阿拉斯加大学安克雷奇分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, Submitted to SIGSPATIAL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12411 2025-06-17 cs.CR cs.CV 79%

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning

Mengyuan Sun, Yu Li, Yuchen Liu, Bo Du, Yunjie Ge

机构 * 1 School of Cyber Science Engineering, Wuhan University 0.3em 2 School of Computer Science, Wuhan University 0.3em 3 Institute for Math \& AI, Wuhan University

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11499 2025-06-16 cs.CL 79%

On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval

Seongbo Jang, Seonghyeon Lee, Dongha Lee, Hwanjo Yu

机构 * Myongji University(明州大学) Kyungpook National University(庆北国立大学) Yonsei University(延世大学) Pohang University of Science and Technology(坡山科学技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments 9 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11526 2025-06-12 cs.AI 79%

RConE: Rough Cone Embedding for Multi-Hop Logical Query Answering on Multi-Modal Knowledge Graphs

Mayank Kharbanda, Rajiv Ratn Shah, Raghava Mutharaju

机构 * IIIT-Delhi(德里印度理工学院)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted in TKDE (June 2025) as regular paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04211 2025-06-12 cs.CL cs.IR 79%

MMREC: LLM Based Multi-Modal Recommender System

Jiahao Tian, Jinman Zhao, Zhenkai Wang, Zhicheng Ding

机构 * Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Toronto(多伦多大学) Columbia University(哥伦比亚大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06912 2025-06-10 cs.CV 79%

Sleep Stage Classification using Multimodal Embedding Fusion from EOG and PSM

Olivier Papillon, Rafik Goubran, James Green, Julien Larivière-Chartier, Caitlin Higginson, Frank Knoefel, Rébecca Robillard

机构 * Bruyère Health Research Institute(布里埃尔健康研究中心) Institute for Mental Health Research at the Royal Ottawa Hospital(皇家渥太华医院心理健康研究所) School of Psychology, University of Ottawa(心理学系,渥太华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Submitted to IEEE MeMeA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19719 2025-06-03 cs.CV 79%

Urban Safety Perception Assessments via Integrating Multimodal Large Language Models with Street View Images

Jiaxin Zhang, Yunqin Li, Tomohiro Fukuda, Bowen Wang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21973 2025-05-29 cs.MM 79%

Towards Structure-aware Model for Multi-modal Knowledge Graph Completion

Linyu Li, Zhi Jin, Yichi Zhang, Dongming Jin, Chengfeng Dou, Yuanpeng He, Xuan Zhang, Haiyan Zhao

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19952 2025-05-27 cs.CV cs.IR 79%

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Rong-Cheng Tu, Wenhao Sun, Hanzhe You, Yingjie Wang, Jiaxing Huang, Li Shen, Dacheng Tao

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算机与数据科学学院) School of Information Science and Technology, University of Science and Technology of China, Hefei, China(中国科学技术大学信息科学与技术学院) Sun Yat-sen University Shenzhen Campus, School of Cyber Science and Technology, Shenzhen, China(中山大学深圳校区计算机科学与技术学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19509 2025-05-27 cs.LG cs.AI 79%

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models

Yifan Jia, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin, Rui Yang, Fenze Feng, Mingcai Chen, Hengyang Lu, Haozhe Wang, Xiaoye Qu, Dongrui Liu, Lizhen Cui, Yuntao Du

机构 * Joint SDU-NTU Centre for Artificial Intelligence Research&School of Software(山东大学与南京工业大学人工智能研究中心及软件学院) University of Science and Technology of China(中国科学技术大学) Shanghai Jiaotong University(上海交通大学) Nanjing University(南京大学) Nanjing University of Posts and Telecommunications(南京邮电大学) Jiangnan University(江南大学) The Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments The source code is available at https://github.com/MLLMKCBENCH/MLLMKC

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11193 2025-05-21 cs.CL 79%

MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model

Jiahao Huo, Yibo Yan, Boren Hu, Yutao Yue, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Tongji University(同济大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by the Main Conference of Empirical Methods in Natural Language Processing (EMNLP) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14014 2025-05-21 cs.CV 79%

EGFormer: Towards Efficient and Generalizable Multimodal Semantic Segmentation

Zelin Zhang, Tao Zhang, KediLI, Xu Zheng

机构 * University of Sydney(悉尼大学) University of Technology Sydney(悉尼技术大学) HKUST(GZ)(香港科技大学(广州))

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13957 2025-05-21 cs.CR cs.CL 79%

Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation

Jiankun Zhang, Shenglai Zeng, Jie Ren, Tianqi Zheng, Hui Liu, Xianfeng Tang, Hui Liu, Yi Chang

机构 * Michigan State University(密歇根州立大学) Jilin University(吉林大学)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13520 2025-05-21 cs.IR cs.AI 79%

Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering

Hessa Alawwad, Usman Naseem, Areej Alhothali, Ali Alkhathlan, Amani Jamal

机构 * Faculty of Computing and Information Technology, King Abdulaziz University(计算机与信息科技学院,国王阿卜杜勒阿齐兹大学) College of Computer and Information Science, Imam Mohammad Ibn Saud Islamic University (IMSIU)(计算机与信息科学学院,伊玛目穆罕默德·本·萨乌德伊斯兰大学) School of Computing, Macquarie University(计算学院,麦考瑞大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 16 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13306 2025-05-20 cs.CV cs.IR 79%

GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval

Chengsong Sun, Weiping Li, Xiang Li, Yuankun Liu, Lianlei Shan

机构 * School of Software and Microelectronics, Peking University(软件与微电子学院,北京大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11815 2025-05-20 cs.CV 79%

UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings

Jiajun Qin, Yuan Pu, Zhuolun He, Seunggeun Kim, David Z. Pan, Bei Yu

机构 * The Chinese University of Hong Kong, China(香港中文大学) ChatEDA Tech(ChatEDA科技) University of Texas at Austin, USA(德克萨斯大学奥斯汀分校)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04960 2025-05-09 cs.IR cs.MM 79%

Learning Item Representations Directly from Multimodal Features for Effective Recommendation

Xin Zhou, Xiaoxiong Zhang, Dusit Niyato, Zhiqi Shen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments Code: https://github.com/enoche/LIRDRec

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18674 2025-05-06 cs.CV cs.LG 79%

Active Data Curation Effectively Distills Large-Scale Multimodal Models

Vishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans, Samuel Albanie, Federico Tombari, Yongqin Xian, Alessio Tonioni, Olivier J. Hénaff

机构 * Google(谷歌) Google DeepMind(谷歌DeepMind) Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学) University of Cambridge(剑桥大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21028 2025-05-01 cs.CR cs.AI cs.LG 79%

Semantic-Aware Contrastive Fine-Tuning: Boosting Multimodal Malware Classification with Discriminative Embeddings

Ivan Montoya Sanchez, Shaswata Mitra, Aritran Piplai, Sudip Mittal

机构 * dept. name of organization (of Aff.)(机构部门名称)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08535 2025-04-29 cs.IR cs.CV cs.LG 79%

Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking

Tianyu Zhu, Myong Chol Jung, Jesse Clark

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Journal ref The ACM Web Conference 2025 (WWW2025) Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10916 2025-04-24 physics.med-ph cs.CV 79%

Embedding Radiomics into Vision Transformers for Multimodal Medical Image Classification

Zhenyu Yang, Haiming Zhu, Rihui Zhang, Haipeng Zhang, Jianliang Wang, Chunhao Wang, Minbin Chen, Fang-Fang Yin

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 27 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13209 2025-04-21 cs.CR cs.AI 79%

On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks

Ting Bi, Chenghang Ye, Zheyu Yang, Ziyi Zhou, Cui Tang, Jun Zhang, Zui Tao, Kailong Wang, Liting Zhou, Yang Yang, Tianlong Yu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07718 2025-04-11 cs.CV 79%

Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval

Zehong Ma, Hao Chen, Wei Zeng, Limin Su, Shiliang Zhang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments TMM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10702 2025-04-11 cs.MM 79%

Retrieval Augmented Verification for Zero-Shot Detection of Multimodal Disinformation

Arka Ujjal Dey, Artemis Llabrés, Ernest Valveny, Dimosthenis Karatzas

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17408 2025-04-11 cs.CL 79%

P-Transformer: A Prompt-based Multimodal Transformer Architecture For Medical Tabular Data

Yucheng Ruan, Xiang Lan, Daniel J. Tan, Hairil Rizal Abdullah, Mengling Feng

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04303 2025-04-08 cs.CL 79%

Graph-Based Multimodal Contrastive Learning for Chart Question Answering

Yue Dai, Soyeon Caren Han, Wei Liu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07202 2025-04-08 cs.AI 79%

A Zero-shot Learning Method Based on Large Language Models for Multi-modal Knowledge Graph Embedding

Bingchen Liu, Jingchen Li, Yuanyuan Fang, Xin Li

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00338 2025-04-02 cs.LG cs.AI cs.MA cs.SI 79%

Agentic Multimodal AI for Hyperpersonalized B2B and B2C Advertising in Competitive Markets: An AI-Driven Competitive Advertising Framework

Sakhinana Sagar Srinivas, Akash Das, Shivam Gupta, Venkataramana Runkana

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15836 2025-03-27 cs.CR cs.AI 79%

Intelligent Code Embedding Framework for High-Precision Ransomware Detection via Multimodal Execution Path Analysis

Levi Gareth, Maximilian Fairbrother, Peregrine Blackwood, Lucasta Underhill, Benedict Ruthermore

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17408 2025-03-25 cs.LG cs.AI 79%

Leveraging OpenFlamingo for Multimodal Embedding Analysis of C2C Car Parts Data

Maisha Binte Rashid, Pablo Rivas

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments The 26th International Conference on Artificial Intelligence (ICAI'24: July 22-25, 2024; Las Vegas, USA)

详情

展开后加载摘要…

URL PDF HTML 收藏