arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-10 至 2026-02-10 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 9 篇

2305.04195 2026-02-10 cs.CV cs.CL 84%

Cross-Modal Retrieval for Motion and Text via DropTriple Loss

通过DropTriple损失实现动与文本的跨模态检索

Sheng Yan, Yang Liu, Haoqiang Wang, Xin Du, Mengyuan Liu, Hong Liu

机构 * School of Artificial Intelligence, Chongqing University of Technology, China(重庆理工大学人工智能学院) College of Computer Science, Sichuan University, China(四川大学计算机学院) Key Laboratory of Machine Perception, Shenzhen Graduate School, Peking University, China(北京大学深圳研究生院机器感知重点实验室)

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 本文提出DropTriple损失,用于提升人类动作与文本之间的跨模态检索性能,实验表明在HumanML3D数据集上实现了较高的检索准确率。

Comments This paper has been accepted by ACM MM Asia 2023 (Best Paper Candidate)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08099 2026-02-10 cs.CV cs.AI 84%

VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

VidVec:解锁视频MLLM嵌入用于视频-文本检索

Issar Tzachor, Dvir Samuel, Rami Ben-Ari

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 VidVec通过利用预训练MLLM的中间层嵌入和校准头部,实现无需训练的视频-文本检索,超越现有方法,达到最佳性能。

Comments Project page: https://iyttor.github.io/VidVec/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07125 2026-02-10 cs.IR cs.AI cs.CV cs.LG 84%

Reasoning-Augmented Representations for Multimodal Retrieval

增强推理的表示用于多模态检索

Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han, Soochahn Lee, Sukanta Ganguly, Yong Jae Lee

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Kookmin University(韩国高丽大学)

专题命中 跨模态检索 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出一种增强推理的多模态检索方法,通过外部化推理和语义密集表示提升检索性能,尤其在知识密集型查询和组合修改请求中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08342 2026-02-10 cs.CV cs.AI 81%

UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Science

UrbanGraphEmbeddings: 学习和评估空间导向的多模态嵌入用于城市科学

Jie Zhang, Xingtong Yu, Yuan Fang, Rudi Stouffs, Zdravko Trivic

机构 * National University of Singapore Department of Architecture Singapore The Chinese University of Hong Kong Dept of Systems Eng. \& Eng. Mgmt. China Singapore Management University School of Computing \& Info. Systems Singapore National University of Singapore Department of Architecture The Chinese University of Hong Kong Singapore Management University

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出UGE框架,通过空间导向的多模态嵌入提升城市理解任务性能,实验显示在图像检索和地理位置排名上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08741 2026-02-10 cs.CL 79%

From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet Understanding

从行到推理:一种增强检索的多模态框架用于电子表格理解

Anmol Gulati, Sahil Sen, Waqar Sarguroh, Kevin Paul

机构 * Commercial Technology and Innovation Office, PricewaterhouseCoopers U.S.(普华永道美国商业技术与创新办公室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 FRTR提出了一种多模态检索增强生成框架,通过分解电子表格为细粒度嵌入并整合多模态信息,提升了对复杂电子表格的推理能力,在基准测试中实现了显著的准确率提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07642 2026-02-10 cs.AI cs.LG 79%

Efficient Table Retrieval and Understanding with Multimodal Large Language Models

基于多模态大语言模型的高效表格检索与理解

Zhuoyan Xu, Haoyang Fang, Boran Han, Bonan Min, Bernie Wang, Cuixiong Hu, Shuai Zhang

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) AWS(亚马逊网络服务)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 TabRAG通过多模态大语言模型实现高效表格检索与理解,显著提升检索召回率和答案准确率。

Comments Published at EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07208 2026-02-10 cs.IR cs.AI 79%

Sequences as Nodes for Contrastive Multimodal Graph Recommendation

序列作为节点的对比多模态图推荐

Bucher Sahyouni, Matthew Vowels, Liqun Chen, Simon Hadfield

机构 * University of Surrey(萨里大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 MuSICRec通过多视图图方法结合协同、序列和多模态信号,提升推荐系统在冷启动和数据稀疏问题上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08700 2026-02-10 cs.CL cs.HC cs.IR 57%

Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search

图像能澄清问题吗?一种研究图像在会话搜索中澄清问题效果的探讨

Clemencia Siro, Zahra Abbasiantaeb, Yifei Yuan, Mohammad Aliannejadi, Maarten de Rijke

机构 * University of Amsterdam(阿姆斯特丹大学) University of Copenhagen(哥本哈根大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 研究探讨图像在会话搜索中澄清问题的效果,发现多模态问题在回答澄清问题时更受青睐,但查询重述任务中效果更平衡,且图像影响因任务类型和用户专业知识而异。

Comments Accepted at CHIIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01317 2026-02-10 cs.HC cs.CY 50%

DietGlance: Dietary Monitoring and Personalized Analysis at a Glance with Knowledge-Empowered AI Assistant

DietGlance: 通过知识增强的AI助手实现饮食监控与个性化分析

Zhihan Jiang, Running Zhao, Lin Lin, Yue Yu, Handi Chen, Xinchen Zhang, Xuhai Xu, Yifang Wang, Xiaojuan Ma, Edith C. H. Ngai

专题命中 跨模态检索 :multimodal(abstract)

AI总结 DietGlance利用知识增强的AI助手,通过眼镜实现日常饮食行为的自动监控和个性化分析,提供营养分析和饮食建议。

Comments 47 pages, 14 figures. Accepted by ACM Transactions on Computing for Healthcare

详情

展开后加载摘要…

URL PDF HTML 收藏