arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-24 至 2025-09-24 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 6 篇

2509.18405 2025-09-24 cs.CV cs.AI 84%

Check Field Detection Agent (CFD-Agent) using Multimodal Large Language and Vision Language Models

Sourav Halder, Jinjun Tong, Xinyu Wu

机构 * U.S. Bank(美国银行)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19087 2025-09-24 cs.CV 79%

Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications

Ganesh Mallya, Yotam Gigi, Dahun Kim, Maxim Neumann, Genady Beryozkin, Tomer Shekel, Anelia Angelova

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19168 2025-09-24 cs.RO 78%

A Multimodal Stochastic Planning Approach for Navigation and Multi-Robot Coordination

Mark Gonzales, Ethan Oh, Joseph Moore

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 8 Pages, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18576 2025-09-24 cs.RO cs.AI 77%

LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA

Zeyi Kang, Liang He, Yanxin Zhang, Zuheng Ming, Kaixing Zhao

机构 * School of Software Northwestern Polytechnical University Xi'an, China(软件学院 西安理工大学中国) Laboratoire L2Tl University Sorbonne Paris Nord Paris, France(L2Tl实验室 索邦巴黎北大学巴黎法国)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18372 2025-09-24 cs.CV 57%

TinyBEV: Cross Modal Knowledge Distillation for Efficient Multi Task Bird's Eye View Perception and Planning

Reeshad Khan, John Gauch

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18609 2025-09-24 cs.RO 50%

PIE: Perception and Interaction Enhanced End-to-End Motion Planning for Autonomous Driving

Chengran Yuan, Zijian Lu, Zhanqi Zhang, Yimin Zhao, Zefan Huang, Shuo Sun, Jiawei Sun, Jiahui Li, Christina Dao Wen Lee, Dongen Li, Marcelo H. Ang

机构 * Department of Mechanical Engineering, National University of Singapore(机械工程系,国立新加坡大学) Department of Civil and Environmental Engineering, National University of Singapore(土木与环境工程系,国立新加坡大学)

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏