arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-06 至 2025-10-06 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 6 篇

2508.00579 2025-10-06 cs.MM cs.IR 79%

MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning

Ziyu Gong, Chengcheng Mai, Yihua Huang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.MM

Comments Comments: Update Title, Author, Abstract, etc

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02580 2025-10-06 cs.AI 79%

V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving

Xuewen Luo, Fengze Yang, Fan Ding, Xiangbo Gao, Shuo Xing, Yang Zhou, Zhengzhong Tu, Chenxi Liu

机构 * University of Utah(犹他大学) Monash University(莫纳什大学) Texas A&M University(德克萨斯农工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06461 2025-10-06 cs.CV 79%

Ranked from Within: Ranking Large Multimodal Models Without Labels

Weijie Tu, Weijian Deng, Dylan Campbell, Yu Yao, Jiyang Zheng, Tom Gedeon, Tongliang Liu

机构 * Australian National University Sydney AI Centre, The University of Sydney Curtin University University of \'OBuda

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICML 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02790 2025-10-06 cs.CV cs.AI cs.CL cs.MM 70%

MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding

Jingyuan Deng, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted to emnlp2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02328 2025-10-06 cs.CL cs.AI cs.MA 62%

AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering

Ziqing Wang, Chengsheng Mao, Xiaole Wen, Yuan Luo, Kaize Ding

机构 * Northwestern University(西北大学) Microsoft(微软公司)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03200 2025-10-06 cs.CV 57%

MonSTeR: a Unified Model for Motion, Scene, Text Retrieval

Luca Collorone, Matteo Gioia, Massimiliano Pappa, Paolo Leoni, Giovanni Ficarra, Or Litany, Indro Spinelli, Fabio Galasso

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) Technion, NVIDIA(技术学院与NVIDIA) WSense

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏