arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2797 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2797 篇

1911.10298 2020-04-03 cs.LG cs.RO stat.ML 78%

CoverNet: Multimodal Behavior Prediction using Trajectory Sets

Tung Phan-Minh, Elena Corina Grigore, Freddy A. Boulton, Oscar Beijbom, Eric M. Wolff

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.09381 2019-06-04 cs.LG stat.ML 78%

Multi-modal Probabilistic Prediction of Interactive Behavior via an Interpretable Model

Yeping Hu, Wei Zhan, Liting Sun, Masayoshi Tomizuka

专题命中 多模态Agent :multi-modal(title,abstract)

Comments accepted by the 2019 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.00499 2019-03-13 cs.CE 78%

Personalized Radiotherapy Design for Glioblastoma: Integrating Mathematical Tumor Models, Multimodal Scans and Bayesian Inference

Jana Lipkova, Panagiotis Angelikopoulos, Stephen Wu, Esther Alberts, Benedikt Wiestler, Christian Diehl, Christine Preibisch, Thomas Pyka, Stephanie Combs, Panagiotis Hadjidoukas, Koen Van Leemput, Petros Koumoutsakos, John S. Lowengrub, Bjoern Menze

专题命中 多模态Agent :multimodal(title,abstract)

Comments Copyright (c) 2019 IEEE. Personal use of this material is permitted. However, permission to use this material for any other purposes must be obtained from the IEEE by sending a request to pubs-permissions@ieee.org. Accepted to IEEE Transactions on Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.00838 2018-10-02 cs.RO 78%

Multimodal Interactive Learning of Primitive Actions

Tuan Do, Nikhil Krishnaswamy, Kyeongmin Rim, James Pustejovsky

专题命中 多模态Agent :multimodal(title,abstract)

Comments Presented at AI-HRI AAAI-FSS, 2018 (arXiv:1809.06606)

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.05481 2018-09-17 cs.DS 78%

Multi-Modal Route Planning in Road and Transit Networks

Daniel Tischner

专题命中 多模态Agent :multi-modal(title,abstract)

Comments Master's thesis, 80 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.01581 2018-09-06 cs.HC 78%

Multimodal Dialogue Management for Multiparty Interaction with Infants

Setareh Nasihati Gilani, David Traum, Arcangelo Merla, Eugenia Hee, Zoey Walker, Barbara Manini, Grady Gallagher, Laura-Ann Petitto

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.03230 2018-09-05 math.PR cs.LG stat.CO stat.ME stat.ML 78%

Does Hamiltonian Monte Carlo mix faster than a random walk on multimodal densities?

Oren Mangoubi, Natesh S. Pillai, Aaron Smith

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.02015 2018-07-27 cs.RO cs.HC 78%

Generative Modeling of Multimodal Multi-Human Behavior

Boris Ivanovic, Edward Schmerling, Karen Leung, Marco Pavone

专题命中 多模态Agent :multimodal(title,abstract)

Comments IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2018 -- 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.10205 2018-05-28 stat.ML cs.LG stat.AP 78%

Multimodal Sentiment Analysis To Explore the Structure of Emotions

Anthony Hu, Seth Flaxman

专题命中 多模态Agent :multimodal(title,abstract)

Comments Accepted as a conference paper at KDD 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.05644 2018-04-17 cs.DS 78%

Multimodal Dynamic Journey Planning

Kalliopi Giannakopoulou, Andreas Paraskevopoulos, Christos Zaroliagis

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.09483 2017-10-27 cs.RO cs.LG 78%

Multimodal Probabilistic Model-Based Planning for Human-Robot Interaction

Edward Schmerling, Karen Leung, Wolf Vollprecht, Marco Pavone

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.06333 2017-08-22 cs.OH 78%

SigViewer: Visualizing Multimodal Signals Stored in XDF (Extensible Data Format) Files

Yida Lin, Clemens Brunner, Paul Sajda, Josef Faller

专题命中 多模态Agent :multimodal(title,abstract)

Comments 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.00470 2017-08-09 stat.ML cs.LG 78%

Learning Multimodal Transition Dynamics for Model-Based Reinforcement Learning

Thomas M. Moerland, Joost Broekens, Catholijn M. Jonker

专题命中 多模态Agent :multimodal(title,abstract)

Comments Scaling Up Reinforcement Learning (SURL) Workshop @ European Machine Learning Conference (ECML)

详情

展开后加载摘要…

URL PDF HTML 收藏
1504.03855 2015-04-17 q-bio.NC physics.med-ph 78%

A versatile clearing agent for multi-modal brain imaging

Irene Costantini, Jean-Pierre Ghobril, Antonino Paolo Di Giovanna, Anna Letizia Allegra Mascaro, Ludovico Silvestri, Marie Caroline Müllenbroich, Leonardo Onofri, Valerio Conti, Francesco Vanzi, Leonardo Sacconi, Renzo Guerrini, Henry Markram, Giulio Iannello, Francesco Saverio Pavone

专题命中 多模态Agent :multi-modal(title);multimodal(abstract)

Comments in Scientific Reports 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1303.4503 2013-03-20 physics.ins-det physics.med-ph 78%

EndoTOFPET-US a Novel Multimodal Tool for Endoscopy and Positron Emission Tomography

Erika Garutti

专题命中 多模态Agent :multimodal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
0809.1074 2009-12-01 math.DS 78%

Multifractal analysis for multimodal maps

Mike Todd

专题命中 多模态Agent :multimodal(title,abstract)

Comments Minor rewrites

详情

展开后加载摘要…

URL PDF HTML 收藏
0711.2531 2009-12-01 q-bio.PE 78%

Multimodal pattern formation in phenotype distributions of sexual populations

Michael Doebeli, Hendrik J. Blok, Olof Leimar, Ulf Dieckmann

专题命中 多模态Agent :multimodal(title,abstract)

Journal ref Proc. R. Soc. B (2007) 274, 347-357

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07246 2024-10-08 cs.CL cs.AI 77%

Propaganda to Hate: A Multimodal Analysis of Arabic Memes with Multi-Agent LLMs

Firoj Alam, Md. Rafiul Biswas, Uzair Shah, Wajdi Zaghouani, Georgios Mikros

专题命中 多模态Agent :multimodal(title,comments);分类 cs.CL、cs.AI

Comments propaganda, hate-speech, disinformation, misinformation, fake news, LLMs, GPT-4, multimodality, multimodal LLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20890 2026-08-24 cs.CV cs.RO 新提交 77%

A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving

面向基于视觉-语言-动作(VLA)的端到端自动驾驶的协同多模态交互

Jingtao Sun, Xiaohai He, Yike Zhang, Dong Huang, Yaonan Wang, Ajmal Mian, Mike Zheng Shou

机构 * National University of Singapore (NUS)(新加坡国立大学) Hunan University(湖南大学) The University of Western Australia (UWA)(西澳大学)

专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文针对现有VLA自动驾驶模型决策不可靠、多模态交互不足的问题,提出含三类核心组件的多模态交互与多轨迹规划系统,实验显示其在安全推理与场景感知上优于现有系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02794 2026-08-21 cs.AI 版本更新 77%

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding

CharTool: 集成工具的视觉推理用于图表理解

Situo Zhang, Yifan Zhang, Zichen Zhu, Da Ma, Lei Pan, Danyang Zhang, Zihan Zhao, Lu Chen, Kai Yu

机构 * X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院X-LANCE实验室) Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室) Suzhou Laboratory(苏州实验室) AISpeech Co., Ltd.(思必驰科技股份有限公司)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 本文提出CharTool,通过集成工具提升多模态大语言模型对图表的理解能力,通过双重数据管道和代理强化学习,在六个图表基准测试中取得显著提升。

Comments Accepted by ACMMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14098 2026-08-13 cs.CV 版本更新 77%

ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization

ForgeryVCR: 通过高效的取证工具在MLLMs中实现视觉中心推理用于图像伪造检测与定位

Youqi Wang, Shen Chen, Haowei Wang, Rongxuan Peng, Taiping Yao, Shunquan Tan, Changsheng Chen, Bin Li, Shouhong Ding

机构 * Shenzhen University(深圳大学) Tencent Youtu Lab(腾讯优图实验室)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 ForgeryVCR通过高效的取证工具实现视觉中心推理,提升图像伪造检测与定位的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09654 2026-08-11 cs.AI 新提交 77%

Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching

无幻觉的GUI定位:基于无回归的布局感知匹配

Yuke Li, Xuehan Hou

机构 * School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 该研究提出无回归框架,通过解耦指令理解与布局感知定位,在ScreenSpot-Pro和Mind2Web数据集上显著提升GUI定位的准确率、成功率及元素选择率,抑制了坐标幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04425 2026-08-11 cs.CL cs.AI cs.CV cs.LG cs.MM 版本更新 77%

UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents

UI-MOPD:用于持续GUI智能体学习的多平台策略蒸馏

Niu Lian, Tongbo Chen, Zhehao Yu, Chengzhen Duan, Fazhan Liu, Hui Liu, Pei Fu, Jian Luan, Heng Qu, Shu-Tao Xia, Jinpeng Wang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Xiaomi(小米) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Zhejiang University(浙江大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 针对构建多平台GUI智能体的挑战,构建高质量数据集Uni-GUI,提出UI-MOPD方法,通过多教师策略蒸馏实现持续学习,动态选教师并转移行为先验,实验证明其能平衡跨平台能力保留与新平台适应。

Comments Technical report. 27 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00066 2026-08-04 cs.CV 新提交 77%

PhysAgent: A Multi-Agent Framework for Reliable Remote Heart Rate Estimation

PhysAgent:用于可靠远程心率估计的多智能体框架

Yehui Yang, Bo Zhao, Junzhe Cao, Hui Ma, Yue Sun, Wenjin Wang, Zitong Yu

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 PhysAgent是一种推理时多智能体候选验证框架,以多个基础rPPG估计器输出为待验证生理假设,结合Qwen3-VL-4B多模态大语言模型推理与确定性融合,提升远程心率估计的稳定性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28595 2026-07-31 cs.CV 新提交 77%

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Beacon:智能体何时及如何执行智能体视觉推理

Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying

机构 * Peking University(北京大学) Kling Team(Kling团队) HKUST(GZ)(香港科技大学(广州)) CUHK(香港中文大学) ZJU(浙江大学) THU(清华大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本研究针对现有智能体视觉推理模型模式适应性有限、工具增益被损害抵消的问题,提出Beacon模型,通过强化学习相关机制提升性能与适应性,在多基准上表现优异。

Comments 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25993 2026-07-29 cs.CV 新提交 77%

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

超越缩放:学习用于超高分辨率遥感的多工具视觉推理

Fengxiang Wang, Jiangnan Huang, Mingshuo Chen, Yueying Li, Yang Shi, Junwei Luo, Haoyu Wang, Yansheng Li, Jing Zhang, Haiyan Zhao, Wenjing Yang

机构 * National University of Defense Technology(国防科技大学) Wuhan University(武汉大学) Tsinghua University(清华大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 针对超高分辨率遥感图像给多模态大语言模型带来的挑战,提出GeoMTVR数据集,结合监督微调与以工具注意力为重点的强化学习算法开发GeoLens,实验证明其在多工具视觉推理方面优于直接推理和单工具放大基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28971 2026-06-30 cs.CV 77%

Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution

自进化智能体图像恢复:通过深思熟虑的规划与直觉执行

Shuang Cui, Fan Ji, Guanglong Sun, Yufei Guo, Xiongxin Tang, Jiangmeng Li, Fanjiang Xu

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) School of Life Sciences, Tsinghua University(清华大学生命科学学院) Intelligent Science & Technology Academy of CASIC(中国科学院 CASIC 智能科学与技术学院)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出SEAR框架,将图像恢复建模为序列决策问题,采用直觉执行器与深思熟虑规划器,结合剪枝感知蒙特卡洛树搜索和自进化情景记忆,解决贪婪搜索和信息利用不足问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26122 2026-06-26 cs.CV 新提交 77%

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents

DocArena:将原始文档转化为可控的训练环境用于文档搜索代理

Jiamian Wang, Ruiyi Zhang, Tong Yu, Jing Shi, Samyadeep Basu, Rajiv Jain, Zhiqiang Tao, Tong Sun

机构 * Rochester Institute of Technology(罗切斯特理工学院) Adobe Research(Adobe研究院)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出DocArena自动化流程,通过多模态文档结构化、推理型QA对构建和质量控制,生成可控训练环境,使基于文本LLM的搜索代理在多模态文档检索和问答中取得最佳性能。

Comments search agent for documents

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22617 2026-06-23 cs.CV 新提交 77%

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

OmniSpace: 自动驾驶多模态大语言模型的高效几何感知

Hao Vo, Phu Loc Nguyen, Khoa Vo, Sieu Tran, Duc Minh Nguyen, Ngo Xuan Cuong, Nghi D. Q. Bui, Anh Nguyen, Duy Minh Ho Nguyen, Ngan Le

机构 * University of Arkansas(阿肯色大学) Google Research, Google(谷歌研究院) University of Liverpool(利物浦大学) Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究所)

专题命中 多模态Agent :MLLM(summary_cn);multimodal(abstract);分类 cs.CV

AI总结 提出OmniSpace,一种即插即用的几何感知范式,通过相机位姿注入器、多视图极线注意力模块和3D几何蒸馏目标,从纯2D观测中提升MLLM的空间推理能力,在多个自动驾驶基准上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09110 2026-06-09 cs.CV 新提交 77%

HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging

HDRAgent: 一种用于多曝光HDR成像的智能体框架

Weiyu Zhou, Tao Hu, Yijian Wang, Xiaogang Xu, Ruixing Wang, Qingsen Yan

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) Shenzhen Research Institute, Northwestern Polytechnical University(西北工业大学深圳研究院) Zhejiang University(浙江大学) Camera Group, DJI(大疆相机部门)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出首个智能体驱动的HDR成像框架HDRAgent,通过细粒度上下文知识匹配、感知-失真反馈机制和智能体引导的生成对齐策略,自适应选择重建策略,减少复杂动态场景中的鬼影和局部伪影。

详情

展开后加载摘要…

URL PDF HTML 收藏