arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2816 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2816 篇

2509.03990 2025-09-09 cs.AI 57%

Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM Agent

Chunlong Wu, Ye Luo, Zhibo Qu, Min Wang

机构 * Tongji University(同济大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00054 2025-09-09 cs.RO cs.AI 57%

Robotic Fire Risk Detection based on Dynamic Knowledge Graph Reasoning: An LLM-Driven Approach with Graph Chain-of-Thought

Haimei Pan, Jiyun Zhang, Qinxi Wei, Xiongnan Jin, Chen Xinkai, Jie Cheng

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments We have decided to withdraw this paper as the work is still undergoing further refinement. To ensure the clarity of the results, we prefer to make additional improvements before resubmission. We appreciate the readers' understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03460 2025-09-09 cs.AI 57%

Multi-Agent Reasoning for Cardiovascular Imaging Phenotype Analysis

Weitong Zhang, Mengyun Qiao, Chengqi Zang, Steven Niederer, Paul M Matthews, Wenjia Bai, Bernhard Kainz

机构 * Department of Computing, Imperial College London, London, UK(帝国理工学院计算机系) Department of Mechanical Engineering, University College London, London, UK(伦敦大学学院机械工程系) Department of Brain Sciences, Imperial College London, London, UK(帝国理工学院脑科学系) Data Science Institute, Imperial College London, London, UK(帝国理工学院数据科学研究所) University of Tokyo, Tokyo, JP(东京大学) National Heart and Lung Institute, Imperial College London, London, UK(帝国理工学院国家心脏和肺研究所) FAU Erlangen-Nürnberg, Erlangen, DE(埃朗根-纽伦堡大学) UK Dementia Research Institute, Imperial College London, London, UK(英国痴呆研究所在伦敦帝国理工学院) Rosalind Franklin Institute, Harwell Science and Innovation Campus, Didcot, UK(罗莎琳德·弗兰克林研究所)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05338 2025-09-09 cs.RO cs.AI 57%

Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks

Atsushi Masumori, Norihiro Maruyama, Itsuki Doi, johnsmith, Hiroki Sato, Takashi Ikegami

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03864 2025-09-09 cs.AI 57%

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety

Zhenyu Pan, Yiting Zhang, Yutong Zhang, Jianshu Zhang, Haozheng Luo, Yuwei Han, Dennis Wu, Hong-Yu Chen, Philip S. Yu, Manling Li, Han Liu

机构 * Northwestern University(西北大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments accepted by the Trustworthy FMs workshop in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15876 2025-09-08 cs.HC cs.AI cs.GR 57%

AI-in-the-loop: The future of biomedical visual analytics applications in the era of AI

Katja Bühler, Thomas Höllt, Thomas Schulz, Pere-Pau Vázquez

机构 * Vienna Research Center for Visual Computing(维也纳视觉计算研究中心) VRVis GmbH(VRVis公司) Delft University of Technology(代尔夫特理工大学) University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所) Universitat Politècnica de Catalunya(巴塞罗那理工大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Computer Graphics & Applications

Journal ref K. Bühler, T. Hollt, T. Schultz and P. Vazquez, "AI-in-The-Loop: The Future of Biomedical Visual Analytics Applications in the Era of AI" in IEEE Computer Graphics and Applications, vol. 45, no. 02, pp. 90-99, March-April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03536 2025-09-05 cs.AI cs.HC 57%

PG-Agent: An Agent Powered by Page Graph

Weizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou, Xiaoxuan Tang, Jiajun Bu, Yong Li, Wei Jiang

机构 * Zhejiang Key Lab of Accessible Perception \& Intelligent Systems, Zhejiang University Hangzhou China Zhejiang University

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Paper accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00218 2025-09-04 cs.RO cs.AI 57%

Embodied AI in Social Spaces: Responsible and Adaptive Robots in Complex Setting -- UKAIRS 2025 (Copy)

Aleksandra Landowska, Aislinn D Gomez Bergin, Ayodeji O. Abioye, Jayati Deshmukh, Andriana Bouadouki, Maria Wheadon, Athina Georgara, Dominic Price, Tuyen Nguyen, Shuang Ao, Lokesh Singh, Yi Long, Raffaele Miele, Joel E. Fischer, Sarvapali D. Ramchurn

机构 * School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院) School of Computing and Communications, The Open University(开放大学计算与通讯学院) School of Electronics and Computer Science, University of Southampton(南安普顿大学电子与计算机科学学院) Responsible AI, University of Southampton(南安普顿大学负责任的人工智能) University of Liverpool(利兹大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18096 2025-09-04 cs.AI 57%

Deep Research Agents: A Systematic Examination And Roadmap

Yuxuan Huang, Yihang Chen, Haozheng Zhang, Kang Li, Huichi Zhou, Meng Fang, Linyi Yang, Xiaoguang Li, Lifeng Shang, Songcen Xu, Jianye Hao, Kun Shao, Jun Wang

机构 * Deep Research Agents(深度研究代理)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02322 2025-09-03 cs.CV 57%

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

Longrong Yang, Zhixiong Zeng, Yufeng Zhong, Jing Huang, Liming Zheng, Lei Chen, Haibo Qiu, Zequn Qin, Lin Ma, Xi Li

机构 * Meituan(美团) Zhejiang University(浙江大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03735 2025-09-03 cs.CV 57%

Multi-Agent System for Comprehensive Soccer Understanding

Jiayuan Rao, Zifeng Li, Haoning Wu, Ya Zhang, Yanfeng Wang, Weidi Xie

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025; Project Page: https://jyrao.github.io/SoccerAgent/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16271 2025-09-03 cs.CV cs.LG 57%

Structuring GUI Elements through Vision Language Models: Towards Action Space Generation

Yi Xu, Yesheng Zhang, Jiajia Liu, Jingdong Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments 10pageV0

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21407 2025-09-03 cs.AI 57%

Graph-Augmented Large Language Model Agents: Current Progress and Future Prospects

Yixin Liu, Guibin Zhang, Kun Wang, Shiyuan Li, Shirui Pan

机构 * Griffith University(格里菲斯大学) National University of Singapore(国立新加坡大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20370 2025-08-29 cs.SE cs.AI 57%

Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought

Lingzhe Zhang, Tong Jia, Kangjin Wang, Weijie Hong, Chiming Duan, Minghua He, Ying Li

机构 * Peking University(北京大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04447 2025-08-27 cs.CV cs.RO 57%

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang, Xinqiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He, Fan Lu, He Wang, Zhizheng Zhang, Li Yi, Wenjun Zeng, Xin Jin

机构 * SJTU(上海交通大学) EIT(欧洲研究所) THU(清华大学) Galbot PKU(北京大学) UIUC(伊利诺伊大学香槟分校) USTC(中国科学技术大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17380 2025-08-26 cs.AI 57%

Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery

Jiaqi Liu, Songning Lai, Pengze Li, Di Yu, Wenjie Zhou, Yiyang Zhou, Peng Xia, Zijun Wang, Xi Chen, Shixiang Tang, Lei Bai, Wanli Ouyang, Mingyu Ding, Huaxiu Yao, Aoran Wang

机构 * UNC–Chapel Hill(北卡罗来纳大学教堂山分校) HKUST (Guangzhou)(香港科技大学(广州)) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Tsinghua University(清华大学) Nankai University(南开大学) UC Santa Cruz(圣塔克鲁兹大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17198 2025-08-26 cs.AI 57%

From reactive to cognitive: brain-inspired spatial intelligence for embodied agents

Shouwei Ruan, Liyuan Wang, Caixin Kang, Qihui Zhu, Songming Liu, Xingxing Wei, Hang Su

机构 * Department of Computer Science and Technology, Institute for AI, BNRist Center, Tsinghua-Bosch Joint ML Center, THBI Lab(计算机科学与技术系、人工智能研究院、BNRist中心、清华-博世联合机器学习中心、THBI实验室) Institute of Artificial Intelligence(人工智能研究院) Department of Psychological and Cognitive Sciences(心理学与认知科学系)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 40 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15164 2025-08-22 cs.CL 57%

ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following

Seungmin Han, Haeun Kwon, Ji-jun Park, Taeyang Yoon

机构 * Dongguk University(东国大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21969 2025-08-22 cs.RO cs.AI 57%

Embodied Long Horizon Manipulation with Closed-loop Code Generation and Incremental Few-shot Adaptation

Yuan Meng, Xiangtong Yao, Haihui Ye, Yirui Zhou, Shengqiang Zhang, Zhenguo Sun, Xukun Li, Zhenshan Bing, Alois Knoll

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments update ICRA 6 page

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10833 2025-08-18 cs.CV 57%

UI-Venus Technical Report: Building High-performance UI Agents with RFT

Zhangxuan Gu, Zhengwen Zeng, Zhenyu Xu, Xingran Zhou, Shuheng Shen, Yunfei Liu, Beitong Zhou, Changhua Meng, Tianyu Xia, Weizhi Chen, Yue Wen, Jingya Dou, Fei Tang, Jinzhen Lin, Yulin Liu, Zhenlin Guo, Yichen Gong, Heng Jia, Changlong Gao, Yuan Guo, Yong Deng, Zhenyu Guo, Liang Chen, Weiqiang Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17593 2025-08-18 cs.HC cs.AI 57%

JELAI: Integrating AI and Learning Analytics in Jupyter Notebooks

Manuel Valle Torre, Thom van der Velden, Marcus Specht, Catharine Oertel

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Accepted for AIED 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09858 2025-08-14 cs.CV 57%

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

Weiqi Li, Zehao Zhang, Liang Lin, Guangrun Wang

专题命中 多模态Agent :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09318 2025-08-14 cs.LO cs.AI 57%

TPTP World Infrastructure for Non-classical Logics

Alexander Steen, Geoff Sutcliffe

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 35 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07466 2025-08-12 cs.AI 57%

Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs

Dom Huh, Prasant Mohapatra

机构 * UC Davis(加州大学戴维斯分校) University of South Florida(佛罗里达州立大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07010 2025-08-12 cs.MM cs.HC cs.MA 57%

Narrative Memory in Machines: Multi-Agent Arc Extraction in Serialized TV

Roberto Balestri, Guglielmo Pescatore

专题命中 多模态Agent :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06342 2025-08-11 cs.CV cs.SI 57%

Street View Sociability: Interpretable Analysis of Urban Social Behavior Across 15 Cities

Kieran Elrod, Katherine Flanigan, Mario Bergés

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00342 2025-08-11 cs.CV 57%

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

Zechuan Li, Hongshan Yu, Yihao Ding, Yan Li, Yong He, Naveed Akhtar

机构 * organization= College of Electrical Information Engineering,Hunan University , city= Changsha , postcode= 410082 , state= Hunan , country= China organization= School of Computing \& Information Systems ,The University of Melbourne , city= Melbourne , postcode= VIC 3053 , state= VIC , country= Australia organization= School of Computer Science,The University of Sydney , city= Sydney , postcode= NSW 2006 , state= NSW , country= Australia organization= School of Artificial Intelligence ,Anhui University , city= Hefei , postcode= 230601 , state= Anhui , country= China

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments This is a submitted version of a paper accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05637 2025-08-11 cs.HC cs.AI 57%

Automated Visualization Makeovers with LLMs

Siddharth Gangwar, David A. Selby, Sebastian J. Vollmer

机构 * University of Kaiserslautern–Landau (RPTU)(凯撒斯劳滕-兰道大学(RPTU)) Department of Data Science and its Applications, German Research Center for Artificial Intelligence (DFKI)(数据科学及其应用系,德国人工智能研究中心(DFKI))

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04389 2025-08-07 cs.AI 57%

GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning

Weitai Kang, Bin Lei, Gaowen Liu, Caiwen Ding, Yan Yan

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Minnesota(明尼苏达大学) Cisco Research(思科研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04239 2025-08-07 cs.CL 57%

DP-GPT4MTS: Dual-Prompt Large Language Model for Textual-Numerical Time Series Forecasting

Chanjuan Liu, Shengzhi Wang, Enqiang Zhu

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏