arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-01 至 2025-08-01 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 5 篇

2507.22896 2025-08-01 cs.HC cs.AI cs.CV cs.RO 84%

iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement

Kohou Wang, ZhaoXiang Liu, Lin Bai, Kun Fan, Xiang Liu, Huan Hu, Kai Wang, Shiguo Lian

机构 * Unicom Data Intelligence(中国联通数据智能研究所) Data Science & Artificial Intelligence Research Institute(数据科学与人工智能研究院) China United Network Communications Group Corporation Limited(中国联合网络通信集团有限公司)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23331 2025-08-01 cs.CV 83%

Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision

Qiang Lu, Waikit Xiu, Xiying Li, Shenyu Hu, Shengbo Sun

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能系统工程学院) Guangdong Provincial Key Laboratory of Intelligent Transportation System(广东省智能交通系统重点实验室)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 11pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23188 2025-08-01 cs.CV 79%

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

Shiyao Yu, Zi-An Wang, Kangning Yin, Zheng Tian, Mingyuan Zhang, Weixin Si, Shihao Zou

机构 * Southern University of Science and Technology(南方科技大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) ShanghaiTech University(上海科技大学) Nanyang Technological University(南洋理工大学) Faculty of Computer Science and Control Engineering, Shenzhen University of Advanced Technology(计算机科学与控制工程学院,深圳大学先进技术学院)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE TMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22938 2025-08-01 cs.CL cs.AI 76%

A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents

Sumit Soman, H. G. Ranjani, Sujoy Roychowdhury, Venkata Dharma Surya Narayana Sastry, Akshat Jain, Pranav Gangrade, Ayaaz Khan

机构 * Ericsson R&D Bangalore Karnataka India(爱立信研发部班加罗尔卡纳塔克邦印度)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL、cs.AI

Comments Accepted for publication at the KDD 2025 Workshop on Structured Knowledge for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23217 2025-08-01 cs.LG cs.AI 57%

Zero-Shot Document Understanding using Pseudo Table of Contents-Guided Retrieval-Augmented Generation

Hyeon Seong Jeong, Sangwoo Jo, Byeong Hyun Yoon, Yoonseok Heo, Haedong Jeong, Taehoon Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏