arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2797 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2797 篇

2602.01983 2026-02-03 cs.AI 74%

Evolving from Tool User to Creator via Training-Free Experience Reuse in Multimodal Reasoning

从工具使用者到创作者的进化:通过无训练经验重用在多模态推理中

Xintian Shen, Jiawei Chen, Lihao Zheng, Hao Ma, Tao Wei, Kun Zhan

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 UCT框架通过无训练经验重用,使智能体从工具使用者转变为工具创造者,提升多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12582 2026-01-21 cond-mat.mtrl-sci cs.AI 74%

Ontology-aligned structuring and reuse of multimodal materials data and workflows towards automatic reproduction

面向多模态材料数据和工作流的本体对齐结构化与重用,以实现自动重现

Sepideh Baghaee Ravari, Abril Azocar Guzman, Sarath Menon, Stefan Sandfeld, Tilmann Hickel, Markus Stricker

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 本文提出一种基于本体驱动和大型语言模型的框架,用于自动提取和结构化多模态材料数据和工作流,以提高计算结果的可重现性和重用性。

Comments 39 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17250 2025-12-22 cs.AI 74%

Accelerating Multi-modal LLM Gaming Performance via Input Prediction and Mishit Correction

通过输入预测和误位修正加速多模态大语言模型游戏性能

Ziyang Lin, Zixuan Sun, Sanhorn Chen, Xiaoyang Chen, Roy Zhao

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

AI总结 通过输入预测和误位修正方法,显著降低多模态大语言模型游戏性能的推理延迟,提升整体控制效果。

Comments UIUC 25 Fall CS 498

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05335 2025-11-04 cs.CV 74%

New multimodal similarity measure for image registration via modeling local functional dependence with linear combination of learned basis functions

Joel Honkamaa, Pekka Marttinen

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments Improved experimental setup

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10255 2025-10-22 cs.CV cs.RO 74%

When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models

Xianzheng Ma, Brandon Smart, Yash Bhalgat, Shuai Chen, Xinghui Li, Jian Ding, Jindong Gu, Dave Zhenyu Chen, Songyou Peng, Jia-Wang Bian, Philip H Torr, Marc Pollefeys, Matthias Nießner, Ian D Reid, Angel X. Chang, Iro Laina, Victor Adrian Prisacariu

机构 * University of Oxford(牛津大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学) Technical University of Munich(慕尼黑技术大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Simon Fraser University(西蒙·弗雷泽大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments 2nd version update to Jun.2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05520 2025-08-27 cs.AI 74%

Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA

Karishma Thakrar, Shreyas Basavatia, Akshay Daftardar

机构 * Georgia Institute of Technology Atlanta, GA, USA(佐治亚理工学院)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03986 2025-08-07 cs.AI 74%

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?

Yuan Xun, Xiaojun Jia, Xinwei Liu, Hua Zhang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Nanyang Technological University(南洋理工大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16940 2025-07-24 cs.CV cs.LG cs.MA 74%

AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation

Nima Fathi, Amar Kumar, Tal Arbel

机构 * Center for Intelligent Machines, McGill University, Montreal, Canada(麦吉尔大学智能机器中心,加拿大蒙特利尔) Mila - Quebec AI institute, Montreal, Canada(魁北克AI研究所)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments 9 pages, 3 figures, International Conference on Medical Image Computing and Computer-Assisted Intervention

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04595 2025-06-06 cs.CV 74%

Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning

Ziqi Jia, Anmin Wang, Xiaoyang Qu, Xiaowen Yang, Jianzong Wang

机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments Accepted by the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21347 2025-05-01 cs.AI cs.HC 74%

IRL Dittos: Embodied Multimodal AI Agent Interactions in Open Spaces

Seonghee Lee, Denae Ford, John Tang, Sasa Junuzovic, Asta Roseway, Ed Cutrell, Kori Inkpen

机构 * Stanford University(斯坦福大学) Microsoft Research(微软研究院)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20028 2025-03-27 cs.AI 74%

OmniNova:A General Multimodal Agent Framework

Pengfei Du

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04730 2025-03-10 cs.CL cs.HC 74%

WinClick: GUI Grounding with Multimodal Large Language Models

Zheng Hui, Yinheng Li, Dan zhao, Tianyi Chen, Colby Banbury, Kazuhito Koishida

专题命中 多模态Agent :multimodal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17288 2024-12-24 cs.RO cs.AI 74%

Multi-Modal Grounded Planning and Efficient Replanning For Learning Embodied Agents with A Few Examples

Taewoong Kim, Byeonghwi Kim, Jonghyun Choi

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

Comments AAAI 2025 (Project page: https://twoongg.github.io/projects/flare/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00627 2024-12-10 cs.HC cs.AI 74%

ARChef: An iOS-Based Augmented Reality Cooking Assistant Powered by Multimodal Gemini LLM

Rithik Vir, Parsa Madinei

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09971 2024-11-18 cs.CV cs.RO 74%

Explanation for Trajectory Planning using Multi-modal Large Language Model for Autonomous Driving

Shota Yamazaki, Chenyu Zhang, Takuya Nanri, Akio Shigekane, Siyuan Wang, Jo Nishiyama, Tao Chu, Kohei Yokosawa

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments Accepted and presented at ECCV 2024 2nd Workshop on Vision-Centric Autonomous Driving (VCAD) on September 30, 2024. 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20409 2024-11-01 cs.CV physics.med-ph 74%

Physics-Regularized Multi-Modal Image Assimilation for Brain Tumor Localization

Michal Balcerak, Tamaz Amiranashvili, Andreas Wagner, Jonas Weidner, Petr Karnakov, Johannes C. Paetzold, Ivan Ezhov, Petros Koumoutsakos, Benedikt Wiestler, Bjoern Menze

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11347 2024-09-18 cs.AI 74%

Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments

Takanori Ugai, Kensho Hara, Shusaku Egami, Ken Fukuda

专题命中 多模态Agent :multimodal(title);分类 cs.AI

Comments 5 pages, 1 figure, 1 table, accepted in Embodied AI 2024 Workshop held in conjunction with CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06978 2024-09-02 cs.CV 74%

Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation

Zhenxin Li, Kailin Li, Shihao Wang, Shiyi Lan, Zhiding Yu, Yishen Ji, Zhiqi Li, Ziyue Zhu, Jan Kautz, Zuxuan Wu, Yu-Gang Jiang, Jose M. Alvarez

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments The 1st place solution of End-to-end Driving at Scale at the CVPR 2024 Autonomous Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06720 2024-08-26 cs.CV cs.LG q-bio.QM 74%

Multimodal Analysis of White Blood Cell Differentiation in Acute Myeloid Leukemia Patients using a β-Variational Autoencoder

Gizem Mert, Ario Sadafi, Raheleh Salehi, Nassir Navab, Carsten Marr

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments Accepted for publication at MICCAI 2024 workshop on AI for Imaging Genomics Learning (AIIG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00535 2024-07-02 cs.CE cs.CV 74%

AI-powered multimodal modeling of personalized hemodynamics in aortic stenosis

Caglar Ozturk, Daniel H. Pak, Luca Rosalia, Debkalpa Goswami, Mary E. Robakowski, Raymond McKay, Christopher T. Nguyen, James S. Duncan, Ellen T. Roche

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments CO and DHP contributed equally to this work. JSD and ETR are corresponding authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05295 2024-06-19 cs.AI cs.FL 74%

Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception

Yunhao Yang, Cyrus Neary, Ufuk Topcu

专题命中 多模态Agent :multimodal(title);分类 cs.AI

Comments Accepted as full paper in AAMAS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16356 2024-04-26 cs.NI cs.AI cs.LG 74%

Integration of Mixture of Experts and Multimodal Generative AI in Internet of Vehicles: A Survey

Minrui Xu, Dusit Niyato, Jiawen Kang, Zehui Xiong, Abbas Jamalipour, Yuguang Fang, Dong In Kim, Xuemin, Shen

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.05821 2023-04-28 cs.LG cs.AI cs.DB 74%

PEg TRAnsfer Workflow recognition challenge report: Does multi-modal data improve recognition?

Arnaud Huaulmé, Kanako Harada, Quang-Minh Nguyen, Bogyu Park, Seungbum Hong, Min-Kook Choi, Michael Peven, Yunshuang Li, Yonghao Long, Qi Dou, Satyadwyoom Kumar, Seenivasan Lalithkumar, Ren Hongliang, Hiroki Matsuzaki, Yuto Ishikawa, Yuriko Harai, Satoshi Kondo, Mamoru Mitsuishi, Pierre Jannin

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

Comments Challenge report doi.org/10.1016/j.cmpb.2023.107561

Journal ref Computer Methods and Programs in Biomedicine, Volume 236, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03996 2022-05-27 cs.LG cs.CV cs.RO 74%

Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers

Ruihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu, Xiaolong Wang

专题命中 多模态Agent :cross-modal(title);分类 cs.CV

Comments Our project page with videos is at https://RchalYang.github.io/LocoTransformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15054 2021-10-29 cs.HC cs.CL cs.CY 74%

Adaptive Multimodal and Multisensory Empathic Technologies for Enhanced Human Communication

Roxana Girju

专题命中 多模态Agent :multimodal(title);分类 cs.CL

Comments 10 pages; This position paper was presented at the Rethinking the Senses: A Workshop on Multisensory Embodied Experiences and Disability Interactions associated with the ACM CHI Conference on Human Factors in Computing Systems, May 2021

Journal ref ACM CHI Conference on Human Factors in Computing Systems, May 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.14156 2021-04-30 cs.RO cs.CV 74%

Radar-based Automotive Localization using Landmarks in a Multimodal Sensor Graph-based Approach

Stefan Jürgens, Niklas Koch, Marc-Michael Meinecke

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Journal ref Proceedings of 21st International Radar Symposium (IRS 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.10384 2021-01-27 cs.RO cs.AI 74%

droidlet: modular, heterogenous, multi-modal agents

Anurag Pratik, Soumith Chintala, Kavya Srinet, Dhiraj Gandhi, Rebecca Qian, Yuxuan Sun, Ryan Drew, Sara Elkafrawy, Anoushka Tiwari, Tucker Hart, Mary Williamson, Abhinav Gupta, Arthur Szlam

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.03199 2020-10-27 cs.CV 74%

Multimodal End-to-End Autonomous Driving

Yi Xiao, Felipe Codevilla, Akhil Gurram, Onay Urfalioglu, Antonio M. López

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments The paper has been accepted by IEEE Transactions on Intelligent Transportation Systems 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.05205 2020-08-11 cs.RO cs.AI cs.MA cs.SY eess.SY 74%

Implicit Multiagent Coordination at Unsignalized Intersections via Multimodal Inference Enabled by Topological Braids

Christoforos Mavrogiannis, Jonathan A. DeCastro, Siddhartha S. Srinivasa

专题命中 多模态Agent :multimodal(title);分类 cs.AI

Comments 16 pages, 13 figures, new experiments, new explanatory figures for intuition and new title

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.03733 2020-02-11 cs.CV 74%

Robust Multimodal Image Registration Using Deep Recurrent Reinforcement Learning

Shanhui Sun, Jing Hu, Mingqing Yao, Jinrong Hu, Xiaodong Yang, Qi Song, Xi Wu

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Journal ref Asian Conference on Computer Vision (ACCV). 2018. 511-526

详情

展开后加载摘要…

URL PDF HTML 收藏