arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2797 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2797 篇

2004.07173 2020-04-16 cs.CV 79%

Bias in Multimodal AI: Testbed for Fair Automatic Recruitment

Alejandro Peña, Ignacio Serna, Aythami Morales, Julian Fierrez

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Journal ref IEEE CVPR Workshop on Fair, Data Efficient and Trusted Computer Vision, Washington, Seattle, USA, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.07093 2020-04-16 cs.LG cs.CL stat.ML 79%

lamBERT: Language and Action Learning Using Multimodal BERT

Kazuki Miyazawa, Tatsuya Aoki, Takato Horii, Takayuki Nagai

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.09746 2020-03-24 cs.AI 79%

Adaptive Informative Path Planning with Multimodal Sensing

Shushman Choudhury, Nate Gruver, Mykel J. Kochenderfer

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments First two authors contributed equally; International Conference on Automated Planning and Scheduling (ICAPS) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.00584 2019-11-05 cs.RO cs.AI cs.LG cs.MA 79%

A Perceived Environment Design using a Multi-Modal Variational Autoencoder for learning Active-Sensing

Timo Korthals, Malte Schilling, Jürgen Leitner

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

Comments Extended Abstract for the IROS 2019 Workshop on Deep Probabilistic Generative Models for Cognitive Architecture in Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.08161 2019-09-19 cs.HC cs.AI cs.RO 79%

Multimodal Continuation-style Architectures for Human-Robot Interaction

Nikhil Krishnaswamy, James Pustejovsky

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Advances in Cognitive Systems Cognitive Vision Workshop (2019), 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.01140 2019-09-04 eess.IV cs.CV 79%

A Tool for Super-Resolving Multimodal Clinical MRI

Mikael Brudfors, Yael Balbastre, Parashkev Nachev, John Ashburner

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.09094 2019-06-24 cs.AI 79%

Hybrid Planning for Dynamic Multimodal Stochastic Shortest Paths

Shushman Choudhury, Mykel J. Kochenderfer

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 20 pages, 5 figures, 5 tables; Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.09949 2019-05-30 cs.RO cs.CV 79%

Scene Induced Multi-Modal Trajectory Forecasting via Planning

Nachiket Deo, Mohan M. Trivedi

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV

Comments ICRA Workshop on Long Term Human Motion Prediction (extended abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.01752 2019-05-07 cs.CV 79%

Understanding urban landuse from the above and ground perspectives: a deep learning, multimodal solution

Shivangi Srivastava, John E. Vargas-Muñoz, Devis Tuia

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Journal ref Remote Sensing of Environment, 228, pages 129 - 143, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01560 2019-05-07 cs.AI cs.RO 79%

Dynamic Real-time Multimodal Routing with Hierarchical Hybrid Planning

Shushman Choudhury, Jacob P. Knickerbocker, Mykel J. Kochenderfer

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 8 figures, Accepted to Intelligent Vehicles (IV) Symposium 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.10763 2018-11-28 cs.CV 79%

Quality-Aware Multimodal Saliency Detection via Deep Reinforcement Learning

Xiao Wang, Tao Sun, Rui Yang, Chenglong Li, Bin Luo, Jin Tang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.09562 2018-07-26 cs.CV 79%

Change Detection between Multimodal Remote Sensing Data Using Siamese CNN

Zhenchao Zhang, George Vosselman, Markus Gerke, Devis Tuia, Michael Ying Yang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.00528 2018-04-03 cs.CV 79%

Multimodal Biometric Authentication Using Choquet Integral and Genetic Algorithm

Anouar Ben Khalifa, Sami Gazzah, Najoua Essoukri Ben Amara

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.02565 2018-02-21 cs.HC cs.AI cs.LG stat.ML 79%

Applying Cooperative Machine Learning to Speed Up the Annotation of Social Signals in Large Multi-modal Corpora

Johannes Wagner, Tobias Baur, Yue Zhang, Michel F. Valstar, Björn Schuller, Elisabeth André

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.04486 2017-10-13 cs.HC cs.CV stat.ML 79%

Multimodal Observation and Interpretation of Subjects Engaged in Problem Solving

Thomas Guntz, Raffaella Balzarini, Dominique Vaufreydaz, James L. Crowley

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Journal ref 1st Workshop on "Behavior, Emotion and Representation: Building Blocks of Interaction'', Oct 2017, Bielefeld, Germany. 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.04376 2016-07-18 cs.RO cs.AI 79%

Intrinsically Motivated Multimodal Structure Learning

Jay Ming Wong, Roderic A. Grupen

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.07806 2016-04-27 cs.AI cs.NE 79%

Using Indirect Encoding of Multiple Brains to Produce Multimodal Behavior

Jacob Schrum, Joel Lehman, Sebastian Risi

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1403.1902 2015-02-04 cs.CV 79%

Quality-based Multimodal Classification Using Tree-Structured Sparsity

Soheil Bahrampour, Asok Ray, Nasser M. Nasrabadi, Kenneth W. Jenkins

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Comments To Appear in 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2014)

Journal ref CVPR 2014, pp. 4114 - 4121

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08183 2026-08-11 cs.RO 新提交 79%

Multi-modal Interactive Control of Robotic Arm based on Offline Large Language Models

基于离线大语言模型的机械臂多模态交互控制

Hanxiao Chen

机构 * University of Tokyo(东京大学)

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(comments)

AI总结 该研究提出“Socratic Models-ChatGLM”算法,基于离线开源大语言模型与PyBullet平台实现机械臂多模态交互控制,可降低成本并解决复杂多步骤机械操作任务。

Comments This research work has been accepted for poster presentation at ICRA 2026 MEI (Multimodal Embodied Interaction in Robots) Workshop. (Here is the workshop-version short paper.)

Journal ref https://ieeexplore.ieee.org/abstract/document/11371765; 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09771 2026-02-11 physics.app-ph 79%

A large scale multi-modal workflow for battery characterization: from concept to implementation

大规模多模态工作流程用于电池表征:从概念到实现

François Cadiou, Cinthya Herrera, Duncan Atkins, Elixabete Ayerbe, Giorgio Baraldi, Stéphanie Belin, Anass Benayad, Didier Blanchard, Federico Capone, Ennio Capria, Isidora Cekic Laskovic, Robert Dominko, Kristina Edström, Ajay Gautam, Lukas Helfen, Antonella Iadecola, Quentin Jacquet, Gregor Kapun, Xinyu Li, Aleksandar Matic, Nataliia Mozhzhukhina, Andrew J Naylor, Poul Norby, Chris O Keefe, Alexandre Ponrouch, Jean Pascal Rueff, Elena Tchernykova, Deyana Tchitchekova, Israel Temprano, Nikita Vostrov, Marnix Wagemaker, Martin Winter, Christian Wölke, Tejs Vegge, Sandrine Lyonnard

专题命中 多模态Agent :multi-modal(title,comments);multimodal(abstract)

AI总结 本文提出了一种大规模多模态工作流程,用于电池表征,通过整合多种技术分析电极材料的演变和电解质成分的影响,展示了标准化流程和多属性元视图的应用。

Comments 32 pages, 6 figures, research article, keywords: Battery, Workflow, Experimental characterization, Multi-modal, Large scale, Correlation, Aging

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09822 2025-07-24 cs.RO cs.SY eess.SY 79%

Active Probing with Multimodal Predictions for Motion Planning

Darshan Gadginmath, Farhad Nawaz, Minjun Sung, Faizan M Tariq, Sangjae Bae, David Isele, Fabio Pasqualetti, Jovin D'sa

机构 * Honda Research Institute, USA(本田研究院(美国)) Department of Mechanical Engineering, University of California Riverside(加州大学河滨分校机械工程系)

专题命中 多模态Agent :multimodal(title,abstract)

Comments To appear at IROS '25. 8 pages. 3 tables. 6 figures. Project page: https://darshangm.github.io/papers/active-probing-multimodal-predictions/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06136 2023-10-11 cs.HC 79%

Predicting Player Engagement in Tom Clancy's The Division 2: A Multimodal Approach via Pixels and Gamepad Actions

Kosmas Pinitas, David Renaudie, Mike Thomsen, Matthew Barthet, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis

专题命中 多模态Agent :multimodal(title,abstract)

Comments 8 pages, accepted for publication and presentation at 2023 25th ACM International Conference on Multimodal Interaction (ICMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12375 2026-07-15 cs.CV cs.AI eess.IV 新提交 79%

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

IQA-T1:基于工具的图像质量评估视觉证据推理

Jinjian Wu, Jiaqi Tang, Wei Wei, Yingying Yan, Jianmin Chen, Botong Geng, Lei Zhang, Qifeng Chen

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机科学学院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 针对开放世界图像质量评估难题,提出IQA-T1框架,通过自主调用工具生成视觉证据增强多模态大语言模型推理,构建Q-Tool数据集,实验表明该框架性能最佳且评估可解释、基于证据。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02542 2026-07-07 cs.AI cs.CV 新提交 79%

iFLYTEK-Embodied-Omni Technical Report

科大讯飞-具身全能技术报告

Yuan Zhang, Jingfei Ni, Guanchen Lu, Shiqi Zhang, Qingshan Xu, Chi Liu, Xin Nie, Wenjie Xu, Lin Gao, Zhiyuan Cheng, Mingxin Zhou, Jiajia Wu, Diyuan Liu, Jia Pan, Chao Ji

机构 * iFLYTEK LindenBot University of Science and Technology of China(中国科学技术大学)

专题命中 多模态Agent :multimodal(abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

AI总结 研究通用具身智能体,提出统一多模态基础模型iFLYTEK-Embodied-Omni,通过共享多模态自注意力联合建模视觉、语言和动作,结合多种数据构建数据集并采用四阶段策略训练,实现脑-小脑协作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29445 2026-06-30 cs.CV cs.AI 79%

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

通过通用关键帧提取桥接VideoQA和视频引导的智能体任务

Sunqi Fan, Qingle Liu, Runqi Yin, Meng-Hao Guo, Shuojin Yang

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出VG-GUIBench基准和TASKER关键帧提取算法,联合考虑任务相关性和场景动态,在VideoQA和视频引导的GUI任务上提升性能。

Comments Accepted by ECCV 2026. Project Page: https://vg-gui-tasker.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26552 2026-06-26 cs.CV cs.AI 新提交 79%

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

感知、判断与进化:基于事后洞察的自优化取证智能体用于AI生成图像检测

Yangjun Wu, Keyu Yan, Yu Liu, Jingren Zhou, Fei Huang, Rong Zhang, Zhou Zhao, Fei Wu

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出ForeAgent框架,采用感知-判断架构融合多视图线索,并引入事后洞察驱动的自优化策略,通过采样-反思-进化范式持续提升检测能力,在多个基准上达到最优性能。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20717 2026-06-23 cs.CV cs.AI cs.CR 新提交 79%

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents

MIRAGE: 针对Web代理的隐蔽视觉提示注入漏洞检测

Xuelong Dai, Jianyu Ma, Boyang Ma, Biwei Yan, Yijun Yang, Yue Zhang

机构 * SDU(山东大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出MIRAGE框架,利用扩散模型在受限区域生成视觉上无害的对抗图像,实现针对多模态大模型Web代理的隐蔽间接提示注入攻击,以检测其视觉漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03005 2026-06-03 cs.CV cs.AI 79%

MUSE: A Unified Agentic Harness for MLLMs

MUSE: 多模态大语言模型的统一智能体框架

Jianglin Lu, Hailing Wang, Xu Ma, Qihua Dong, Mingyuan Zhang, Yizhou Wang, Yun Fu

机构 * Northeastern University(东北大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出MUSE框架,通过可组合模块(任务表示、视觉处理、感知工具、结构化解析、确定性验证和验证器引导修复)提升冻结多模态大语言模型性能,无需重新训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30639 2026-06-01 cs.CV cs.AI cs.RO 79%

PInVerify: An Offline Embodied Benchmark for Active Instance Verification

PInVerify:面向主动实例验证的离线具身基准

Yuhang Jiang

机构 * University of Trento(特伦托大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出主动实例验证任务,构建离线具身基准PInVerify,通过多视角导航和细粒度属性匹配评估具身智能体,并基于多模态大语言模型建立基线。

Comments Accepted as a poster at the Foundation Models Meet Embodied Agents (FMEA) Workshop, CVPR 2026. 44 pages including appendix. Code: https://github.com/Avalon-S/PInVerify

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24023 2026-01-01 cs.CV cs.AI 79%

RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations

RSAgent: 通过多轮工具调用学习推理与行动进行文本引导的分割

Xingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li, Lingyi Hong, Mingxi Chen, Kaixun Jiang, Jiyuan Fu, Wenqiang Zhang

机构 * Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University(上海智能信息处理关键实验室,计算机科学与人工智能学院,复旦大学) Artificial Intelligence, Fudan University(人工智能,复旦大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 RSAgent通过多轮工具调用实现文本引导分割的推理与行动,采用两阶段框架提升分割性能,达到领域内和领域外基准的最先进水平。

详情

展开后加载摘要…

URL PDF HTML 收藏