arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3475 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3475 篇

2505.03135 2026-01-09 cs.AI 79%

Beyond Retrieval: Improving Evidence Quality for LLM-based Multimodal Fact-Checking

超越检索:提升基于大语言模型的多模态事实核查的证据质量

Haoran Ou, Gelei Deng, Xingshuo Han, Jie Zhang, Han Qiu, Shangwei Guo, Tianwei Zhang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出Aletheia框架,通过改进证据检索策略提升多模态事实核查的证据质量,实验证明其在虚假信息检测中的高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18867 2026-01-07 cs.AI 79%

Topological Perspectives on Optimal Multimodal Embedding Spaces

拓扑视角下的最优多模态嵌入空间

Abdul Aziz A. B, A. B Abdul Rahim

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文通过拓扑数据分析比较CLIP和CLOOB的嵌入空间,揭示其模态差距驱动因素和维度坍缩的影响,为多模态模型优化提供新视角。

Comments This manuscript contains substantive technical inaccuracies and an incomplete treatment of the stated topic. Subsequent developments and a reassessment of the problem indicate that the scope and framing of the work do not adequately reflect the current state of research, and the analysis is therefore incomplete and outdated

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01718 2026-01-06 cs.AI 79%

Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications

Yuan3.0 Flash:面向企业应用的开源多模态大语言模型

YuanLab. ai, :, Shawn Wu, Sean Wang, Louie Li, Darcy Chen, Allen Wang, Jiangang Luo, Xudong Zhao, Joseph Shen, Gawain Ma, Jasper Jia, Marcus Mao, Claire Wang, Hunter He, Carol Wang, Zera Zhang, Jason Wang, Chonly Shen, Leo Zhang, Logan Chen, Qasim Meng, James Gong, Danied Zhao, Penn Zheng, Owen Zhu, Tong Yu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 Yuan3.0 Flash通过RAPO算法提升企业任务性能,实现高效多模态大语言模型开源

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10393 2026-01-06 cs.SE cs.AI 79%

Cross-modal Retrieval Models for Stripped Binary Analysis

用于剥离二进制分析的跨模态检索模型

Guoqiang Chen, Lingyun Ying, Ziyang Song, Daguang Liu, Qiang Wang, Zhiqi Wang, Li Hu, Shaoyin Cheng, Weiming Zhang, Nenghai Yu

机构 * University of Science and Technology of China(中国科学技术大学) QI-ANXIN Technology Research Institute(QI-ANXIN技术研究所)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出BinSeek模型,通过两阶段跨模态检索框架实现对剥离二进制代码的高效检索,显著提升二进制代码与自然语言描述的相关性检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00623 2026-01-05 cs.AI 79%

DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations

DA-DPO:面向减少多模态大语言模型幻觉的高效难度感知偏好优化

Longtian Qiu, Shan Ning, Chuyu Zhang, Jiaxuan Sun, Xuming He

机构 * ShanghaiTech University(上海科技大学) Lingang Laboratory(灵冈实验室) Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心)

专题命中 跨模态检索 :MLLM(title);multimodal(abstract);分类 cs.AI

AI总结 DA-DPO通过难度感知机制优化多模态大语言模型的偏好学习,有效减少幻觉并提升模型鲁棒性与泛化能力。

Comments Accepted by TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13754 2025-12-30 cs.CV 79%

Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval

跨模态全模式细粒度对齐用于文本到图像人物检索

Hao Yin, Xin Man, Feiyu Chen, Jie Shao, Heng Tao Shen

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(深圳先进研究所,电子科学与技术大学) University of Electronic Science and Technology of China(电子科学与技术大学) Sichuan Artificial Intelligence Research Institute(四川人工智能研究院)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出FMFA框架,通过显式细粒度对齐和隐式关系推理实现文本到图像人物检索的高精度匹配。

Comments accepted by ACM Transactions on Multimedia Computing Communications and Applications in December 2025

Journal ref ACM Transactions on Multimedia Computing Communications and Applications, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21698 2025-12-29 cs.CR cs.MM eess.IV 79%

Raster Domain Text Steganography: A Unified Framework for Multimodal Secure Embedding

位图域文本隐写术:一种多模态安全嵌入的统一框架

A V Uday Kiran Kandala

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出了一种基于位图域的多模态安全嵌入框架,通过字形扰动和像素计数实现文本、图像、音频和视频等异构数据的隐蔽传输。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11712 2025-12-23 cs.AI 79%

Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization

通过理论一致的对称多模态偏好优化缓解幻觉

Wenqi Liu, Xuemeng Song, Jiaxi Li, Yinwei Wei, Na Zheng, Jianhua Yin, Liqiang Nie

机构 * Shandong University(山东大学) Southern University of Science and Technology(南方科技大学) University of Georgia(佐治亚大学) National University of Singapore(新加坡国立大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 SymMPO通过理论一致的对称多模态偏好优化方法,有效缓解多模态大语言模型中的幻觉问题。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17194 2025-12-22 cs.AI 79%

MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation

MMRAG-RFT: 两阶段强化微调用于可解释的多模态检索增强生成

Shengwei Zhao, Jingwen Yao, Sitong Wei, Linhai Xu, Yuying Liu, Dong Zhang, Zhiqiang Tian, Shaoyi Du

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 MMRAG-RFT通过两阶段强化微调提升多模态检索增强生成的可解释性,实现更清晰的推理逻辑和更优的生成效果。

Comments This paper was accepted to AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16802 2025-12-19 cs.CL 79%

Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology

多模态检索增强生成在生物医学领域中的增强策略探索:一项评估糖生物学问答的案例研究

Primož Kocbek, Azra Frkatović-Hodžić, Dora Lalić, Vivian Hui, Gordan Lauc, Gregor Štiglic

机构 * University of Maribor, Faculty of Health Sciences(莫拉维亚大学健康科学学院) University of Ljubljana, Medical Factory(卢布尔雅那大学医疗工厂) Genos Ltd(基因公司) Center for Smart Health, School of Nursing The Hong Kong Polytechnic University(智能健康中心护理学院香港理工大学) University of Zagreb, Faculty of Pharmacy(扎格雷布大学药学院) Usher Institute University of Edinburgh(埃德蒙顿大学usher研究所)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

AI总结 本文研究了多模态检索增强生成在生物医学领域中的增强策略,通过实验发现多模态转换和视觉检索在不同模型中均能提升问答准确率。

Comments Will be published in IEEE BigData 2025 proceedings. Contains 10 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23316 2025-12-16 cs.CV 79%

C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

C3-OWD: 一种用于开放世界检测的课程跨模态对比学习框架

Siheng Wang, Zhengdao Li, Yanshu Li, Canran Xiao, Haibo Zhan, Zhengtao Yao, Xuzhi Zhang, Jiale Kang, Linshan Li, Weiming Liu, Zhikang Dong, Jifeng Shen, Junhao Dong, Qiang Sun, Piotr Koniusz

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 C3-OWD提出一种课程跨模态对比学习框架,通过预训练和视觉-语言对齐提升目标检测的鲁棒性和泛化能力。

Comments one of the authors doesn't agree any more

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10419 2025-12-12 cs.CV 79%

TransLocNet: Cross-Modal Attention for Aerial-Ground Vehicle Localization with Contrastive Learning

TransLocNet: 跨模态注意力用于航空-地面车辆定位的对比学习

Phu Pham, Damon Conover, Aniket Bera

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 TransLocNet通过跨模态注意力和对比学习实现航空-地面车辆定位,显著提升定位精度和鲁棒性。

Comments 8 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05463 2025-12-08 cs.CY cs.AI 79%

Modelling Business Agreements in the Multimodal Transportation Domain through Ontological Smart Contracts

通过本体智能合约建模多模态交通运输领域的商业协议

Mario Scrocca, Marco Comerio, Alessio Carenini, Irene Celino

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出通过本体智能合约建模多模式交通运输领域的商业协议,展示其在拼车场景中的应用及优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05036 2025-12-03 cs.CL 79%

From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models

从词向量到多模态嵌入:大型语言模型的技术、应用与未来方向

Charles Zhang, Benji Peng, Xintian Sun, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Yichao Zhang, Xinyuan Song, Cheng Fei, Caitlyn Heqi Yin, Lawrence KQ Yan, Hongyang He, Tianyang Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Simon Fraser University(西蒙弗雷泽大学) Kyoto University(京都大学) National Taiwan Normal University(台湾师范大学) Purdue University(普渡大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校) Emory University(埃默里大学) Cornell University(康奈尔大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The Hong Kong University of Science(香港科学大学) University of Liverpool(利物浦大学) University of Warwick(沃里克大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 本文综述了从词向量到多模态嵌入的发展,探讨了大型语言模型的技术、应用及未来方向。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22997 2025-12-01 cs.CV cs.RO 79%

MrGS: Multi-modal Radiance Fields with 3D Gaussian Splatting for RGB-Thermal Novel View Synthesis

MrGS: 多模态辐射场与3D高斯点云融合用于RGB-热成像新视角合成

Minseong Kweon, Janghyun Kim, Ukcheol Shin, Jinsun Park

机构 * Minnesota Robotics Institute (MnRI), University of Minnesota, Twin Cities(明尼苏达大学罗学院(MnRI)、明尼苏达大学双城分校) Department of Information Convergence Engineering (Artificial Intelligence Major), Pusan National University(信息融合工程系(人工智能专业),釜山国立大学) Department of Energy Engineering, Korea Institute of Energy Technology (KENTECH)(能源工程系,韩国能源技术研究所(KENTECH)) School of Computer Science and Engineering, Pusan National University(计算机科学与工程学院,釜山国立大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 MrGS通过多模态辐射场与3D高斯点云融合,实现RGB和热成像新视角合成,利用物理原理建模热传导和辐射现象,提升重建精度与效率。

Comments Accepted at Thermal Infrared in Robotics (TIRO) Workshop, ICRA 2025 (Best Poster Award)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19380 2025-11-25 cs.CV 79%

UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval

UISearch: 基于图的多模态企业UI截图检索

Maroun Ayli, Youssef Bakouny, Tushar Sharma, Nader Jalloul, Hani Seifeddine, Rima Kilany

机构 * Center For Computer Science(计算机科学中心) Saint Joseph University of Beirut(贝鲁特圣约瑟夫大学) Faculty of Computer Science(计算机科学学院) Dalhousie University(达尔豪斯大学) Murex

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 UISearch通过基于图的结构嵌入与语义检索结合,实现多模态企业UI截图检索,提升检索准确率与效率。

Comments 12 pages, 2 figures, 3 algorithms, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18983 2025-11-25 cs.CV 79%

UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection

UMCL: 单模生成多模对比学习用于跨压缩率深度伪造检测

Ching-Yi Lai, Chih-Yu Jian, Pei-Cheng Chuang, Chia-Ming Lee, Chih-Chung Hsu, Chiou-Ting Hsu, Chia-Wen Lin

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 UMCL通过单模生成多模对比学习,提升跨压缩率深度伪造检测的鲁棒性和准确性。

Comments 24-page manuscript accepted to IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07419 2025-11-24 cs.IR cs.MM 79%

Breaking the Curse of Knowledge: Towards Effective Multimodal Recommendation using Knowledge Soft Integration

突破知识诅咒:通过知识软整合实现有效的多模态推荐

Kai Ouyang, Chen Tang, Zenghao Chai, Wenhao Zheng, Xiangjin Xie, Xuanji Xiao, Zhi Wang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出KSI框架,通过知识软整合解决多模态推荐中的知识诅咒问题,提升推荐个性化效果。

Comments Accepted to IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13189 2025-11-18 cs.CV cs.IR 79%

Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework

Diego Ortego, Marlon Rodríguez, Mario Almagro, Kunal Dahiya, David Jiménez, Juan C. SanMiguel

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments To appear at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08181 2025-11-18 cs.IR cs.AI 79%

MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System

Seung Hwan Cho, Yujin Yang, Danik Baeck, Minjoo Kim, Young-Min Kim, Heejung Lee, Sangjin Park

机构 * Department of Industrial Data Engineering, Hanyang University, Republic of Korea(工业数据工程系,翰阳大学) School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea(跨学科工业研究学院,翰阳大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 2 figures, Accepted at RDGENAI at CIKM 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17714 2025-11-11 cs.CV 79%

F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model

Hanbo Bi, Zhiqiang Yuan, Zexi Jia, Jiapei Zhang, Chongyang Li, Peixiang Luo, Ying Deng, Xiaoyue Duan, Jinchao Zhang

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17494 2025-11-11 eess.IV cs.CV 79%

Enhancing Multimodal Medical Image Classification using Cross-Graph Modal Contrastive Learning

Jun-En Ding, Chien-Chin Hsu, Chi-Hsiang Chu, Shuqiang Wang, Feng Liu

机构 * Department of Systems Engineering, Stevens Institute of Technology, Hoboken, New Jersey, USA(系统工程系,史蒂文斯理工学院) Department of Nuclear Medicine, Kaohsiung Chang Gung Memorial Hospital, Kaohsiung, Taiwan(高雄长庚纪念医院核医学部) Institute of Statistics, National University of Kaohsiung, Kaohsiung, Taiwan(国立高雄大学统计研究所) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China(深圳先进技术研究院,中国科学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03070 2025-11-11 cs.LG cs.MM 79%

FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization

Manh Duong Nguyen, Trung Thanh Nguyen, Huy Hieu Pham, Trong Nghia Hoang, Phi Le Nguyen, Thanh Trung Huynh

机构 * Hanoi University of Science and Technology(河内科学技术大学) Nagoya University(名古屋大学) Washington State University(华盛顿州立大学) Swiss Federal Institute of Technology Lausanne(洛桑联邦理工学院)

专题命中 跨模态检索 :cross-modal(title);multi-modal(abstract);分类 cs.MM

Comments The 22nd International Symposium on Network Computing and Applications (NCA 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06268 2025-11-11 cs.CV cs.CY 79%

LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval

Jian Zhang, Junyi Guo, Junyi Yuan, Huanda Lu, Yanlin Zhou, Fangyu Wu, Qiufeng Wang, Dongming Lu

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) NingboTech University(宁波科技学院) Dunhuang Academy(敦煌研究院) Zhejiang University(浙江大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01892 2025-11-05 cs.LG cs.CL 79%

Retrieval-Augmented Multimodal Depression Detection

Ruibo Hou, Shiyu Teng, Jiaqing Liu, Shurong Chai, Yinhao Li, Lanfen Lin, Yen-Wei Chen

机构 * College of Information Science and Engineering, Ritsumeikan University(信息科学与工程学院,立命馆大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted in IEEE EMBC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00903 2025-11-04 cs.CL 79%

ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval

Ahmed Masry, Megh Thakkar, Patrice Bechard, Sathwik Tejaswi Madhusudhan, Rabiul Awal, Shambhavi Mishra, Akshay Kalkunte Suresh, Srivatsava Daruru, Enamul Hoque, Spandana Gella, Torsten Scholak, Sai Rajeswar

机构 * ServiceNow York University(约克大学) MILA - Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) École de technologie supérieure(卓越工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27350 2025-11-03 cs.CV 79%

RzenEmbed: Towards Comprehensive Multimodal Retrieval

Weijian Jian, Yajun Zhang, Dawei Liang, Chunyu Xie, Yixiao He, Dawei Leng, Yuhui Yin

机构 * AI Research(360人工智能研究院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23224 2025-10-28 cs.CV cs.IR 79%

Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment

Hongyi Wang, Zhengjie Zhu, Jiabo Ma, Fang Wang, Yue Shi, Bo Luo, Jili Wang, Qiuyu Cai, Xiuming Zhang, Yen-Wei Chen, Lanfen Lin, Hao Chen

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学计算机科学与工程系) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Department of Radiology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology(华中科技大学同济医学院附属同济医院放射科) Department of Pathology, Sir Run Run Shaw Hospital, School of Medicine, Zhejiang University(浙江大学医学院附属邵氏医院病理科) Department of Pathology, The Central Hospital of Wuhan, Tongji Medical College, Huazhong University of Science and Technology(华中科技大学同济医学院附属武汉中心医院病理科) Department of Pathology, The First Affiliated Hospital, School of Medicine, Zhejiang University(浙江大学医学院附属第一医院病理科) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学化学与生物工程系) Division of Life Science, The Hong Kong University of Science and Technology(香港科学与技术大学生命科学系)

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22880 2025-10-28 cs.LG cs.AI 79%

Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data

Duong M. Nguyen, Trong Nghia Hoang, Thanh Trung Huynh, Quoc Viet Hung Nguyen, Phi Le Nguyen

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Washington State University(华盛顿州立大学) VinUniversity(文大学) Griffin University(格里芬大学) Hanoi University of Science and Technology(河内科学技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05715 2025-10-28 cs.IR cs.MM 79%

From ID-based to ID-free: Rethinking ID Effectiveness in Multimodal Collaborative Filtering Recommendation

Guohao Li, Li Jing, Jia Wu, Xuefei Li, Kai Zhu, Yue He

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments We identified that our current approach achieves its reported performance only under specific data conditions, and its robustness is weaker than we initially expected

详情

展开后加载摘要…

URL PDF HTML 收藏