arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-14 至 2025-10-14 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 7 篇

2510.10828 2025-10-14 cs.IR cs.AI 79%

VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering

Zhenghan Tai, Hanwei Wu, Qingchen Hu, Jijun Chi, Hailin He, Lei Ding, Tung Sum Thomas Kwok, Bohuai Xiao, Yuchen Hua, Suyuchen Wang, Peng Lu, Muzhi Li, Yihong Wu, Liheng Ma, Jerry Huang, Jiayi Zhang, Gonghao Zhang, Chaolong Jiang, Jingrui Tian, Sicheng Lyu, Zeyu Li, Boyu Han, Fengran Mo, Xinyue Yu, Yufei Cui, Ling Zhou, Xinyu Wang

机构 * University of Toronto(多伦多大学) McMaster University(麦马斯特大学) McGill University(麦吉尔大学) University of Manitoba(曼尼托巴大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Montreal(蒙特利尔大学) Mila CUHK(香港中文大学) HKUST(GZ)(香港理工大学(广州)) Nanyang Technological University(南洋理工大学) Stanford University(斯坦福大学) CG Matrix Technology Limited(CG矩阵科技有限公司)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21524 2025-10-14 cs.CV cs.LG stat.ML 70%

Learning Shared Representations from Unpaired Data

Amitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri Shaham

机构 * Department of Computer Science Bar-Ilan University(巴伊兰大学计算机科学系) Electrical and Computer Engineering Technion(技术学院电子与计算机工程系)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10426 2025-10-14 cs.CV cs.AI 62%

Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs

Suyang Xi, Chenxi Yang, Hong Ding, Yiqing Ni, Catherine C. Liu, Yunhao Liu, Chengqi Zhang

机构 * Emory University(埃默里大学) University of Electronic Science and Technology of China(电子科技大学) University of Illinois Chicago(伊利诺伊大学香槟分校) The Hong Kong Polytechnic University(香港理工大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11204 2025-10-14 cs.CV 57%

Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos

Rohit Gupta, Anirban Roy, Claire Christensen, Sujeong Kim, Sarah Gerard, Madeline Cincebeaux, Ajay Divakaran, Todd Grindal, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) SRI International(SRI国际)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Published at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10787 2025-10-14 cs.CL 57%

Review of Inference-Time Scaling Strategies: Reasoning, Search and RAG

Zhichao Wang, Cheng Wan, Dong Nie

机构 * Inflection AI Georgia Institute of Technology(佐治亚理工学院) ChatAlpha AI

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05970 2025-10-14 cs.CV 57%

Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval

Haiwen Li, Delong Liu, Zhaohui Hou, Zhicheng Zhao, Fei Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments This paper was originally submitted to ACM MM 2025 on April 12, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10655 2025-10-14 q-bio.OT 56%

Isotropy and Geometry of Pretrained Protein LMs

Sheikh Azizul Hakim, Kowshic Roy, M Saifur Rahman

专题命中 跨模态检索 :multi-modal(abstract,comments)

Comments Published in the Proceedings of the ICML 2025 Workshop on Multi-modal Foun- dation Models and Large Language Models for Life Sciences, Vancouver, Canada. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏