arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2790 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2790 篇

2507.07306 2025-07-11 cs.AI cs.CL eess.AS 82%

ViDove: A Translation Agent System with Multimodal Context and Memory-Augmented Reasoning

Yichen Lu, Wei Dai, Jiaen Liu, Ching Wing Kwok, Zongheng Wu, Xudong Xiao, Ao Sun, Sheng Fu, Jianyuan Zhan, Yian Wang, Takatomo Saito, Sicheng Lai

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24382 2025-06-02 cs.RO eess.SP 82%

MagicGripper: A Multimodal Sensor-Integrated Gripper for Contact-Rich Robotic Manipulation

Wen Fan, Haoran Li, Dandan Zhang

机构 * Department of Bioengineering, Imperial College London(帝国理工学院伦敦分校生物工程系) School of Robotics, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学机器人学院)

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract)

Comments 19 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16865 2025-02-25 cs.IR 82%

Multimodal Search in Chemical Documents and Reactions

Ayush Kumar Shah, Abhisek Dey, Leo Luo, Bryan Amador, Patrick Philippy, Ming Zhong, Siru Ouyang, David Mark Friday, David Bianchi, Nick Jackson, Richard Zanibbi, Jiawei Han

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract)

Comments 4 pages, 2 figures, SIGIR 2025 Demonstration Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14528 2025-02-21 math.OC 82%

Dynamic Preference-based Multi-modal Trip Planning of Public Transport and Shared Mobility

Yimeng Zhang, Oded Cats, Shadi Sharif Azadeh

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00252 2025-02-18 cs.AI cs.CL cs.CV cs.MA 82%

Towards Rationality in Language and Multimodal Agents: A Survey

Bowen Jiang, Yangxinyu Xie, Xiaomeng Wang, Yuan Yuan, Zhuoqun Hao, Xinyi Bai, Weijie J. Su, Camillo J. Taylor, Tanwi Mallick

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper has been accepted to the NAACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14394 2025-02-13 cs.AI cs.CL cs.CV 82%

A Multimodal Automated Interpretability Agent

Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram, Evan Hernandez, Jacob Andreas, Antonio Torralba

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 25 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12574 2025-01-24 cs.AI cs.CL cs.CV cs.LG 82%

MuMA-ToM: Multi-modal Multi-Agent Theory of Mind

Haojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin, Leyla Isik, Yen-Ling Kuo, Tianmin Shu

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI-25 (Oral). Project website: https://scai.cs.jhu.edu/projects/MuMA-ToM/ Code: https://github.com/SCAI-JHU/MuMA-ToM

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11051 2025-01-22 cs.CV cs.AI cs.CL cs.RO 82%

FLAME: Learning to Navigate with Multimodal LLM in Urban Environments

Yunzhe Xu, Yiyuan Pan, Zhe Liu, Hesheng Wang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to AAAI 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11974 2024-12-18 cs.RO cs.AI cs.CL cs.CV 82%

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Qi Sun, Pengfei Hong, Tej Deep Pala, Vernon Toh, U-Xuan Tan, Deepanway Ghosal, Soujanya Poria

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments https://github.com/declare-lab/Emma-X, https://huggingface.co/declare-lab/Emma-X

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08442 2024-12-12 cs.LG 82%

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, Devon Hjelm, Zhe Gan, Zsolt Kira, Alexander Toshev

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07904 2024-12-10 cs.LG 82%

Grounding Multimodal Large Language Models in Actions

Andrew Szot, Bogdan Mazoure, Harsh Agrawal, Devon Hjelm, Zsolt Kira, Alexander Toshev

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10603 2024-11-19 cs.RO cs.SY eess.SY 82%

A Novel MLLM-based Approach for Autonomous Driving in Different Weather Conditions

Sonda Fourati, Wael Jaafar, Noura Baccar

专题命中 多模态Agent :MLLM(title,abstract);multi-modal(abstract)

Comments 9 pages, 6 figures; Submitted to IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21480 2024-10-30 cs.LG cs.AI cs.CL cs.CV 82%

AiSciVision: A Framework for Specializing Large Multimodal Models in Scientific Image Classification

Brendan Hogan, Anmol Kabra, Felipe Siqueira Pacheco, Laura Greenstreet, Joshua Fan, Aaron Ferber, Marta Ummus, Alecsander Brito, Olivia Graham, Lillian Aoki, Drew Harvell, Alex Flecker, Carla Gomes

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14277 2024-09-24 cs.AI cs.CL cs.CV cs.RO 82%

Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models

Yew Ken Chia, Qi Sun, Lidong Bing, Soujanya Poria

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06327 2024-08-13 cs.AI cs.CL cs.CV 82%

VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, Jiadai Sun, Xinyue Yang, Yu Yang, Zehan Qi, Shuntian Yao, Xueqiao Sun, Siyi Cheng, Qinkai Zheng, Hao Yu, Hanchen Zhang, Wenyi Hong, Ming Ding, Lihang Pan, Xiaotao Gu, Aohan Zeng, Zhengxiao Du, Chan Hee Song, Yu Su, Yuxiao Dong, Jie Tang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14972 2024-08-07 cs.AI cs.CL cs.MA cs.MM 82%

A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning

Changmeng Zheng, Dayong Liang, Wengyu Zhang, Xiao-Yong Wei, Tat-Seng Chua, Qing Li

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted by ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02121 2024-08-06 physics.flu-dyn physics.app-ph physics.comp-ph 82%

Non-invasive imaging assisted CFD simulation of 4D multi-modal fluid flow using In-situ adaptor

Vaishali Sharma, Arpit Kumar, Snehlata Shakya, Mayank Goswami

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract)

Comments 11 Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18035 2024-07-26 cs.CV cs.AI cs.CL 82%

RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models

Haoyu Chen, Wenbo Li, Jinjin Gu, Jingjing Ren, Sixiang Chen, Tian Ye, Renjing Pei, Kaiwen Zhou, Fenglong Song, Lei Zhu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01587 2024-06-05 cs.RO 82%

PlanAgent: A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning

Yupeng Zheng, Zebin Xing, Qichao Zhang, Bu Jin, Pengfei Li, Yuhang Zheng, Zhongpu Xia, Kun Zhan, Xianpeng Lang, Yaran Chen, Dongbin Zhao

专题命中 多模态Agent :multi-modal(title,abstract);MLLM(abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18358 2024-05-29 cs.CL cs.AI cs.CV cs.LG 82%

MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning

Somnath Kumar, Yash Gadhia, Tanuja Ganu, Akshay Nambi

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16829 2024-05-27 cs.CV cs.AI cs.CL 82%

Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials

Ye Fang, Zeyi Sun, Tong Wu, Jiaqi Wang, Ziwei Liu, Gordon Wetzstein, Dahua Lin

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://sunzey.github.io/Make-it-Real/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18137 2024-05-27 cs.RO cs.AI cs.CL cs.CV cs.LG 82%

DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning

Jianxiong Li, Jinliang Zheng, Yinan Zheng, Liyuan Mao, Xiao Hu, Sijie Cheng, Haoyi Niu, Jihao Liu, Yu Liu, Jingjing Liu, Ya-Qin Zhang, Xianyuan Zhan

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11640 2024-05-21 cs.AI cs.CL cs.CV 82%

Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning

Zishan Gu, Fenglin Liu, Changchang Yin, Ping Zhang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04950 2024-05-09 cs.CV cs.AI cs.CL 82%

VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context

Yunxin Li, Baotian Hu, Haoyuan Shi, Wei Wang, Longyue Wang, Min Zhang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 17 pages; Accepted by ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08268 2023-10-12 cs.RO cs.AI cs.CL cs.LG cs.SD eess.AS 82%

Chat with the Environment: Interactive Multimodal Perception Using Large Language Models

Xufeng Zhao, Mengdi Li, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments IROS2023, Detroit. See the project website at https://matcha-agent.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.07363 2020-11-17 cs.IR 82%

RecTen: A Recursive Hierarchical Low Rank Tensor Factorization Method to Discover Hierarchical Patterns in Multi-modal Data

Risul Islam, Md Omar Faruk Rokon, Evangelos E. Papalexakis, Michalis Faloutsos

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract)

Comments 9 pages, 9 figures, 1 table, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14593 2026-07-20 cs.HC cs.AI cs.CL 版本更新 82%

Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction

记忆驱动的自我表露与关系转折点:人机交互的纵向多模态研究

Ryuichi Sumida, Mao Saeki, Masaki Eguchi, Sadahiro Yoshikawa, Koji Inoue, Tatsuya Kawahara, Yoichi Matsuyama

机构 * Graduate School of Informatics, Kyoto University(京都大学信息学研究科) Equmenopolis, Inc.(Equmenopolis公司) Waseda University(早稻田大学) School of Informatics, Kyoto University(京都大学信息学系)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究通过纵向多模态研究,探讨对话式人工智能系统中交互如何发展为关系。核心方法是让参与者对五个关系构建要素评分,发现对话质量影响当下愉悦感,感知记忆受关系制约,关系有崩溃和激增等转折点,揭示了人机关系建立的方式。

Comments 15 pages, 3 figures. Accepted to ICMI 2026 (International Conference on Multimodal Interaction), October 5-9, 2026, Napoli, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02629 2025-08-07 cs.RO cs.AI cs.CL 82%

HyCodePolicy: Hybrid Language Controllers for Multimodal Monitoring and Decision in Embodied Agents

Yibin Liu, Zhixuan Liang, Zanxin Chen, Tianxing Chen, Mengkang Hu, Wanxi Dong, Congsheng Xu, Zhaoming Han, Yusen Qin, Yao Mu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI;multi-modal(comments)

Comments Accepted to ICCV 2025 Workshop on Multi-Modal Reasoning for Agentic Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16170 2023-12-27 cs.CV cs.AI cs.RO 82%

EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen, Wenwei Zhang, Kai Chen, Tianfan Xue, Xihui Liu, Cewu Lu, Dahua Lin, Jiangmiao Pang

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments A multi-modal, ego-centric 3D perception dataset and benchmark for holistic 3D scene understanding. Project page: http://tai-wang.github.io/embodiedscan

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11844 2026-08-05 cs.CV 版本更新 81%

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

超越单摄像头:体育视频理解中的智能多视角推理

Kerui Chen, Jinglu Wang, Xiaoyi Zhang, Yan Lu

机构 * Zhejiang University(浙江大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态Agent :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV

AI总结 针对体育视频多视角理解缺乏评估基准及MLLMs难以利用多视角信息的问题,引入SportMV - Bench基准,分析瓶颈所在,并提出SportMV - Agent框架,通过迭代循环实现主动视角选择等,相比最强MLLM基线有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏