arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2816 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2816 篇

2511.19033 2025-11-25 cs.CV 57%

ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay

ReEXplore: 通过上下文化回顾经验回放提升具身探索的MLLMs

Gengyuan Zhang, Mingcong Ding, Jingpei Wu, Ruotong Liao, Volker Tresp

机构 * LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心) TU Munich(慕尼黑技术大学)

专题命中 多模态Agent :MLLM(abstract);分类 cs.CV

AI总结 ReEXplore通过回顾经验回放和分层前沿选择,提升MLLMs在具身探索中的效率和性能。

Comments 8 main pages plus 13 pages Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15253 2025-11-25 cs.HC cs.AI 57%

PresentCoach: Dual-Agent Presentation Coaching through Exemplars and Interactive Feedback

PresentCoach: 通过典范和交互反馈的双智能体演示指导

Sirui Chen, Jinsong Zhou, Xinli Xu, Xiaoyu Yang, Litao Guo, Ying-Cong Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 PresentCoach通过双智能体系统,结合典范和交互反馈,提供高效的演示技能指导,提升学习者在教育和专业领域的表现能力。

Comments 13pages,6figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21651 2025-11-24 cs.AI 57%

Can AI Perceive Physical Danger and Intervene?

AI能否感知物理危险并干预?

Abhishek Jindal, Dmitry Kalashnikov, R. Alex Hofer, Oscar Chang, Divya Garikapati, Anirudha Majumdar, Pierre Sermanet, Vikas Sindhwani

机构 * Google DeepMind Robotics(谷歌深Mind机器人技术)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出了一种用于评估具身体验AI系统物理安全性的基准测试方法,通过生成逼真图像和视频来测试模型对安全风险的理解和干预能力,并开发了训练后范式以提升模型的安全推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23463 2025-11-24 cs.CV 57%

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

OpenDriveVLA: 向端到端自动驾驶迈进的大型视觉语言动作模型

Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma, Volker Tresp, Alois Knoll

机构 * Technical University of Munich(慕尼黑技术大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 OpenDriveVLA基于开源大语言模型,通过多模态输入和分层视觉语言对齐,实现端到端自动驾驶中的高精度轨迹规划和驾驶任务回答。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15456 2025-11-20 cs.AI q-fin.GN 57%

Know Your Intent: An Autonomous Multi-Perspective LLM Agent Framework for DeFi User Transaction Intent Mining

了解您的意图:一种自主多视角大语言模型代理框架用于DeFi用户交易意图挖掘

Qian'ang Mao, Yuxuan Zhang, Jiaman Chen, Wenjun Zhou, Jiaqi Yan

机构 * Nanjing University(南京大学) The University of Tennessee, Knoxville(田纳西大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出TIM框架,通过多视角LLM代理系统挖掘DeFi用户交易意图,提升意图推断的准确性和可验证性。

Comments Written in 2025 Q1

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23596 2025-11-20 cs.AI 57%

Agent-SAMA: State-Aware Mobile Assistant

Linqiang Guo, Wei Liu, Yi Wen Heng, Tse-Hsun, Chen, Yang Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to AAAI-26 (Main Technical Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04614 2025-11-18 cs.AI 57%

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

Yuyang Wanyan, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Jiabo Ye, Yutong Kou, Ming Yan, Fei Huang, Xiaoshan Yang, Weiming Dong, Changsheng Xu

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences, China(自动化研究所,中国科学院,中国) School of Artificial Intelligence, University of Chinese Academy of Sciences, China(中国科学院大学人工智能学院,中国) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00108 2025-11-17 cs.LG cs.AI cs.RO 57%

Pelican-VL 1.0: A Foundation Brain Model for Embodied Intelligence

Yi Zhang, Che Liu, Xiancong Ren, Hanchu Ni, Shuai Zhang, Zeyuan Ding, Jiayu Hu, Hanzhe Shan, Zhenwei Niu, Zhaoyang Liu, Shuang Liu, Yue Zhao, Junbo Qi, Qinfan Zhang, Dengjie Li, Yidong Wang, Jiachen Luo, Yong Dai, Zenglin Xu, Bin Shen, Qifan Wang, Jian Tang, Xiaozhu Ju

机构 * WFM System Group(WFM系统组) Beijing Innovation Center of Humanoid Robotics (X-Humanoid)(北京人形机器人创新中心(X-Humanoid))

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09727 2025-11-14 cs.RO cs.AI cs.LG 57%

Baby Sophia: A Developmental Approach to Self-Exploration through Self-Touch and Hand Regard

Stelios Zarifis, Ioannis Chalkiadakis, Artemis Chardouveli, Vasiliki Moutzouri, Aggelos Sotirchos, Katerina Papadimitriou, Panagiotis Filntisis, Niki Efthymiou, Petros Maragos, Katerina Pastra

机构 * Robotics Institute, Athena Research Center(机器人研究所,阿提卡研究中心) HERON – Hellenic Robotics Center of Excellence(HERON – 希腊机器人 excellence 中心) Institute for Language and Speech Processing, Athena Research Center(语言和语音处理研究所,阿提卡研究中心)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 5 pages, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08732 2025-11-13 cs.RO cs.AI cs.LG 57%

Intuitive Programming, Adaptive Task Planning, and Dynamic Role Allocation in Human-Robot Collaboration

Marta Lagomarsino, Elena Merlo, Andrea Pupa, Timo Birr, Franziska Krebs, Cristian Secchi, Tamim Asfour, Arash Ajoudani

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Published in the Annual Review of Control, Robotics, and Autonomous Systems, Volume 9; copyright 2026 the author(s), CC BY 4.0, https://www.annualreviews.org

Journal ref Annual Review of Control, Robotics, and Autonomous Systems (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08521 2025-11-12 cs.CV 57%

UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist

Zhengyang Liang, Daoan Zhang, Huichi Zhou, Rui Huang, Bobo Li, Yuechen Zhang, Shengqiong Wu, Xiaohan Wang, Jiebo Luo, Lizi Liao, Hao Fei

机构 * Singapore Management University(新加坡管理大学) University of Rochester(罗切斯特大学) University College London(伦敦大学学院) National University of Singapore(新加坡国立大学) The Chinese University of Hong Kong(香港中文大学) Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Technical Report. 24 figures, 37 pages. Website: https://univa.online/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24563 2025-11-12 cs.CV 57%

OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents

Hongrui Jia, Jitong Liao, Xi Zhang, Haiyang Xu, Tianbao Xie, Chaoya Jiang, Ming Yan, Si Liu, Wei Ye, Fei Huang

机构 * Peking University(北京大学) Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) Beijing Zhongguancun Academy(北京中关村学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11075 2025-11-12 cs.AI 57%

Emergence of Goal-Directed Behaviors via Active Inference with Self-Prior

Dongmin Kim, Hoshinori Kanazawa, Naoto Yoshida, Yasuo Kuniyoshi

机构 * Graduate School of Information Science and Technology(信息科学与技术研究生院) The University of Tokyo(东京大学) Graduate School of Informatics(信息学研究生院) Kyoto University(京都大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 23 pages, 8 figures, Code is available at https://github.com/kim135797531/self-prior

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06260 2025-11-11 cs.GT cs.AI cs.SY eess.SY 57%

LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling

Hanlin Sun, Jiayang Li

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05639 2025-11-11 cs.CL 57%

ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?

Haoxin Wang, Xianhan Peng, Xucheng Huang, Yizhe Huang, Ming Gong, Chenghan Yang, Yang Liu, Ling Jiang

机构 * Xiaoduo AI Lab(小多人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03265 2025-11-10 cs.LG cs.AI 57%

Cognitive Edge Computing: A Comprehensive Survey on Optimizing Large Models and AI Agents for Pervasive Deployment

Xubin Wang, Qing Li, Weijia Jia

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02532 2025-11-05 cs.AI cs.LG 57%

Agentic AI for Mobile Network RAN Management and Optimization

Jorge Pellejero, Luis A. Hernández Gómez, Luis Mendo Tomás, Zoraida Frias Barroso

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04633 2025-11-05 cs.NE cs.AI cs.LG q-bio.NC 57%

The Physical Basis of Prediction: World Model Formation in Neural Organoids via an LLM-Generated Curriculum

Brennen Hill

机构 * Department of Computer Science University of Wisconsin-Madison(计算机科学系 威斯康星大学麦迪逊分校)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Published in the proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Scaling Environments for Agents (SEA). Additionally accepted for presentation in NeurIPS 2025 Workshop: Embodied World Models for Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00096 2025-11-04 cs.MA cs.AI cs.CY 57%

Urban-MAS: Human-Centered Urban Prediction with LLM-Based Multi-Agent System

Shangyu Lou

机构 * University of California, Santa Barbara \& San Diego State University California USA University of California, Santa Barbara \& San Diego State University

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to The 3rd ACM SIGSPATIAL International Workshop on Advances in Urban AI (UrbanAI'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00933 2025-11-04 cs.RO cs.CV 57%

Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation

Xiangyu Shi, Zerui Li, Yanyuan Qiao, Qi Wu

机构 * Australian Institute for Machine Learning, the University of Adelaide(澳大利亚机器学习研究所、阿德莱德大学) CREATE Lab, Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院CREATE实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13178 2025-11-03 cs.CR cs.AI cs.RO 57%

SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents

Sheng Yin, Xianghe Pang, Yuanzhuo Ding, Menglan Chen, Yutong Bi, Yichen Xiong, Wenhao Huang, Zhen Xiang, Jing Shao, Siheng Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Shanghai Jiao Tong University(上海交通大学) Department of Computer Science(计算机科学系) University of Georgia(佐治亚大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态Agent :image-text(abstract);分类 cs.AI

Comments 28 pages, 19 tables, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26022 2025-10-31 eess.IV cs.CV 57%

Groupwise Registration with Physics-Informed Test-Time Adaptation on Multi-parametric Cardiac MRI

Xinqi Li, Yi Zhang, Li-Ting Huang, Hsiao-Huang Chang, Thoralf Niendorf, Min-Chi Ku, Qian Tao, Hsin-Jung Yang

机构 * Biomedical Imaging Research Institute, Cedars-Sinai Medical Center, United States Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine in the Helmholtz Association (MDC), Germany Department of Imaging Physics, Delft University of Technology, Netherlands TUM School of Computation, Information Technology, Technische Universität München, Germany Department of Surgery, Taipei Veterans General Hospital, Taiwan

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17197 2025-10-31 cs.LG cs.AI eess.SP 57%

SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing

Junlong Ke, Qiying Hu, Shenghai Yuan, Yuecong Xu, Jianfei Yang

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院) MARS Lab, Nanyang Technological University, Singapore(南洋理工大学MARS实验室)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21566 2025-10-28 cs.MA cs.CL 57%

ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem

Fangwen Wu, Zheng Wu, Jihong Wang, Yunku Chen, Ruiguang Pei, Heyuan Huang, Xin Liao, Xingyu Lou, Huarong Deng, Zhihui Fu, Weiwen Liu, Zhuosheng Zhang, Weinan Zhang, Jun Wang

机构 * Shanghai Jiao Tong University(上海交通大学) OPPO

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21669 2025-10-28 cs.AI 57%

SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents

Wanxin Tian, Shijie Zhang, Kevin Zhang, Xiaowei Chi, Chunkai Fan, Junyu Lu, Yulin Luo, Qiang Zhou, Yiming Zhao, Ning Liu, Siyu Lin, Zhiyuan Qin, Xiaozhu Ju, Shanghang Zhang, Jian Tang

机构 * Beijing Innovation Center of Humanoid Robotics(人形机器人创新中心) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12733 2025-10-27 cs.RO cs.AI cs.LG 57%

HYPE: Hybrid Planning with Ego Proposal-Conditioned Predictions

Hang Yu, Julian Jordan, Julian Schmidt, Silvan Lindner, Alessandro Canevaro, Wilhelm Stork

机构 * Mercedes-Benz AG, Research & Development(梅赛德斯-奔驰集团,研发部) Karlsruhe Institute of Technology, ITIV(卡尔斯鲁厄理工学院,ITIV) University of Tübingen(图宾根大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to IEEE ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16924 2025-10-21 cs.CL 57%

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?

Zhihui Yang, Yupei Wang, Kaijie Mo, Zhe Zhao, Renfen Hu

机构 * Beijing Normal University(北京师范大学) Tencent AI Lab(腾讯AI实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (Findings). This version corrects a redundant sentence in the Results section that appeared in the camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16572 2025-10-21 cs.AI cs.MA 57%

Ripple Effect Protocol: Coordinating Agent Populations

Ayush Chopra, Aman Sharma, Feroz Ahmad, Luca Muscariello, Vijoy Pandey, Ramesh Raskar

机构 * Massachusetts Institute of Technology(麻省理工学院) Project Iceberg Cisco(思科)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16457 2025-10-21 cs.CV cs.RO 57%

NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation

Peiran Xu, Xicheng Gong, Yadong MU

机构 * Peking University(北京大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14194 2025-10-17 cs.AI 57%

Implementation of AI in Precision Medicine

Göktuğ Bender, Samer Faraj, Anand Bhardwaj

机构 * Desautels Faculty of Management(德萨尔斯管理学院) McGill University(麦吉尔大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to SMASH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏