arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-07 至 2026-01-07 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 5 篇

2601.02365 2026-01-07 cs.IR cs.AI cs.CL cs.LG 82%

FUSE : Failure-aware Usage of Subagent Evidence for MultiModal Search and Recommendation

FUSE : 多模态搜索与推荐中子代理证据的故障感知使用

Tushar Vatsa, Vibha Belavadi, Priya Shanmugasundaram, Suhas Suresha, Dewang Sultania

机构 * Adobe Inc.(Adobe公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 FUSE通过上下文压缩策略提升多模态搜索与推荐性能,实现93.3%的意图准确率和99.4%的召回率。

Comments ICDM MMSR 2025: Workshop on Multimodal Search and Recommendations

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15201 2026-01-07 cs.CV cs.MM 82%

Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval

迈向食品图像到食谱检索的无偏跨模态表示学习

Qing Wang, Chong-Wah Ngo, Ee-Peng Lim

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM

AI总结 本文通过因果理论解决食品图像到食谱检索中的偏见问题,提出因果干预方法和多标签分类器,提升检索性能。

Comments Code link: https://github.com/GZWQ/Towards-Unbiased-Cross-Modal-Representation-Learning-for-Food-Image-to-Recipe-Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18867 2026-01-07 cs.AI 79%

Topological Perspectives on Optimal Multimodal Embedding Spaces

拓扑视角下的最优多模态嵌入空间

Abdul Aziz A. B, A. B Abdul Rahim

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文通过拓扑数据分析比较CLIP和CLOOB的嵌入空间,揭示其模态差距驱动因素和维度坍缩的影响,为多模态模型优化提供新视角。

Comments This manuscript contains substantive technical inaccuracies and an incomplete treatment of the stated topic. Subsequent developments and a reassessment of the problem indicate that the scope and framing of the work do not adequately reflect the current state of research, and the analysis is therefore incomplete and outdated

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21121 2026-01-07 cs.IR cs.AI 57%

Beyond Patch Aggregation: 3-Pass Pyramid Indexing for Vision-Enhanced Document Retrieval

超越补丁聚合:面向视觉增强文档检索的三阶段金字塔索引

Anup Roy, Rishabh Gyanendra Upadhyay, Animesh Rameshbhai Panara, Robin Mills, Aidan Millar

机构 * Inception AI Mubadala

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 VisionRAG是一种无OCR、模型无关的多模态检索系统,通过三阶段金字塔索引提升文档检索效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02807 2026-01-07 cs.IR cs.LG 50%

COFFEE: COdesign Framework for Feature Enriched Embeddings in Ads-Ranking Systems

COFFEE:广告排序系统中特征增强嵌入的协同设计框架

Sohini Roychowdhury, Doris Wang, Qian Ge, Joy Mu, Srihari Reddy

机构 * Meta Platforms, Inc.(Meta平台公司)

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 COFFEE框架通过整合多样化事件源、延长用户历史和多模态嵌入,提升广告排序系统中用户-广告表示的准确性和规模效应。

Comments 4 pages, 5 figures, 1 table

Journal ref WSDM, Web and Graph Workshop, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏