arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2797 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2797 篇

2606.13190 2026-06-12 cs.RO cs.HC 新提交 71%

Multi-Modal Multi-Agent Robotic Cognitive Alignment enabled by Non-Invasive Consumer Brain Computer Interfaces: A Proof of Concept Exploration

基于非侵入式消费级脑机接口的多模态多智能体机器人认知对齐:概念验证探索

Nataliya Kosmyna, Liz Jenkins, Anoop K. Sinha

机构 * GOOGLE(谷歌) Paradigms of Intelligence(智能范式) Cambridge, MA, United States(马萨诸塞州剑桥市,美国) Mountain View, CA, United States(加利福尼亚州山景城,美国)

专题命中 多模态Agent :multi-modal(title)

AI总结 提出一种框架,利用消费级脑机接口监测脑电信号,在高认知负荷时延迟智能体通信,实现认知对齐的多智能体交互,初步验证了实时信号处理、大语言模型与机器人结合的可行性。

Comments 19 pages, 9 figures, for associated video, see https://youtu.be/0Tav-G87XGs

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01774 2026-05-12 cs.RO cs.SY eess.SY 71%

MOBIUS: A Multi-Modal Bipedal Robot that can Walk, Crawl, Climb, and Roll

MOBIUS:一种能够行走、爬行、攀爬和滚动的多模态双足机器人

Alexander Schperberg, Yusuke Tanaka, Stefano Di Cairano, Dennis Hong

机构 * Mitsubishi Electric Research Laboratories(三菱电机研究实验室) Robotic Systems Lab(机器人系统实验室) Robotics and Mechanisms Laboratory(机器人与机构实验室) Department of Mechanical and Aerospace Engineering, University of California, Los Angeles(加州大学洛杉矶分校机械与航空航天工程系)

专题命中 多模态Agent :multi-modal(title)

AI总结 MOBIUS机器人通过四条肢体实现多种运动模式切换,结合强化学习与混合规划架构,实现动态攀爬与负载支撑,拓展了移动操作与抓取能力。

Comments Paper is accepted at the Robotics: Science and Systems conference, held in Sydney, Australia, July 13th-17th, 2026. Alexander Schperberg and Yusuke Tanaka are co-first authors. Both were at the Robotics and Mechanisms Laboratory (RoMeLa) at UCLA when the work started, and are now with Mitsubishi Electric Research Laboratories and ETH Zurich (RSL) respectively

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21410 2026-03-24 cs.RO 71%

Bayesian Active Object Recognition and 6D Pose Estimation from Multimodal Contact Sensing

基于多模态接触传感的贝叶斯主动物体识别与6D位姿估计

Haodong Zheng, Gabriele M. Caddeo, Andrei C. Jalba, Wijnand A. IJsselsteijn, Lorenzo Natale, Raymond H. Cuijpers

机构 * Eindhoven University of Technology(埃因霍温理工大学) Italian Institute of Technology(意大利理工学院) Humanoid Sensing and Perception Group(人形感知与感知小组)

专题命中 多模态Agent :multimodal(title)

AI总结 本文提出一种结合触觉与自由空间约束的贝叶斯框架,用于联合物体识别和6D位姿估计,通过定制粒子滤波器提升推理效率,并利用推理结果指导主动探索,实验表明触觉信息显著提升了识别和位姿估计的准确性与稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08260 2026-03-10 cs.RO 71%

Seed2Scale: A Self-Evolving Data Engine for Embodied AI via Small to Large Model Synergy and Multimodal Evaluation

Seed2Scale: 一种通过小到大模型协同和多模态评估的自我进化数据引擎用于具身AI

Cong Tai, Zhaoyu Zheng, Haixu Long, Hansheng Wu, Zhengbin Long, Haodong Xiang, Rong Shi, Zhuo Cui, Shizhuang Zhang, Gang Qiu, He Wang, Ruifeng Li, Biao Liu, Zhenzhe Sun, Tao Shen

机构 * ZTE Corporation(中兴通讯公司)

专题命中 多模态Agent :multimodal(title)

AI总结 Seed2Scale通过小到大模型协同和多模态评估,实现具身AI的自我进化数据引擎,显著提升性能并提供可扩展的开发路径

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16895 2026-02-20 cs.HC 71%

Connecting the Dots: Surfacing Structure in Documents through AI-Generated Cross-Modal Links

连接点:通过AI生成的跨模态链接揭示文档结构

Alyssa Hwang, Hita Kambhamettu, Yue Yang, Ajay Patel, Joseph Chee Chang, Andrew Head

专题命中 多模态Agent :cross-modal(title)

AI总结 本文提出通过AI生成跨模态链接的框架,帮助用户更高效地理解和整合复杂文档中的信息。

Comments 40 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10648 2026-02-12 cs.HC 71%

Generative Muscle Stimulation: Providing Users with Physical Assistance by Constraining Multimodal-AI with Embodied Knowledge

生成性肌肉刺激:通过将多模态AI与具身知识相结合,为用户提供物理帮助

Yun Ho, Romain Nith, Peili Jiang, Steven He, Bruno Felalaga, Shan-Yuan Teng, Rhea Seeralan, Pedro Lopes

专题命中 多模态Agent :multimodal(title)

AI总结 本研究提出通过结合多模态AI与具身知识生成肌肉刺激指令,实现更通用的物理辅助系统。

Comments 22 pages, 29 figures

Journal ref Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10809 2026-01-28 cs.CR cs.LG 71%

MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents

针对代理的恶意图像补丁:多模态操作系统代理劫持

Lukas Aichberger, Alasdair Paren, Guohao Li, Philip Torr, Yarin Gal, Adel Bibi

机构 * Johannes Kepler University Linz(约翰内斯·开普勒大学林茨) University of Oxford(牛津大学)

专题命中 多模态Agent :multimodal(title)

AI总结 本文提出恶意图像补丁(MIPs)攻击,通过篡改屏幕区域使多模态OS代理执行有害操作,揭示了OS代理的安全漏洞。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03630 2025-12-04 cs.RO 71%

Multimodal Control of Manipulators: Coupling Kinematics and Vision for Self-Driving Laboratory Operations

多模态操作臂控制:结合运动学与视觉的自我驱动实验室操作

Shifa Sulaiman, Amarnath H, Simon Bogh, Naresh Marturi

机构 * 1 Department of Electronics Systems, Aalborg University, Denmark 3 Extreme Robotics Laboratory, School of Metallurgy \& Materials, University of Birmingham, Birmingham, United Kingdom

专题命中 多模态Agent :multimodal(title)

AI总结 本文提出三种基于雅可比方法的运动规划方案,结合RRT*算法和螺旋理论,用于优化冗余机械臂与耦合手指夹具的轨迹规划与逆解计算,以提升实验室自动化操作的效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20904 2025-11-27 cs.IR 71%

Generating Querying Code from Text for Multi-Modal Electronic Health Record

从文本生成多模态电子健康记录的查询代码

Mengliang ZHang

专题命中 多模态Agent :multi-modal(title)

AI总结 本文提出TQGen-EHRQuery框架,通过整合表格和文本数据,提升多模态电子健康记录查询的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24413 2025-09-30 cs.RO cs.HC 71%

DynaMIC: Dynamic Multimodal In-Context Learning Enabled Embodied Robot Counterfactual Resistance Ability

Tianqiang Yan, Ziqiao Lin, Sicheng Wang, Tianwei Zhang, Zhenglong Sun

机构 * Faculty of Information Technology, Monash University(墨尔本大学信息技术学院) School of Science and Engineering, the Chinese University of Hong Kong-Shenzhen(香港中文大学(深圳)科学与工程学院) Shenzhen Institute of Artificial Intelligence and Robotics for Society, the Chinese University of Hong Kong-Shenzhen(深圳人工智能与机器人研究院)

专题命中 多模态Agent :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00586 2025-08-14 cs.RO cs.LG 71%

ParkDiffusion: Heterogeneous Multi-Agent Multi-Modal Trajectory Prediction for Automated Parking using Diffusion Models

Jiarong Wei, Niclas Vödisch, Anna Rehr, Christian Feist, Abhinav Valada

机构 * Department of Computer Science, University of Freiburg(弗赖堡大学计算机科学系) CARIAD SE(CARIAD公司)

专题命中 多模态Agent :multi-modal(title)

Comments IROS 2025 Camera-Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07894 2025-08-08 math.OC 71%

Complexity Analysis of a Bicriteria Directed Multimodal Transportation Network Design Problem

Dominik Leib, Susanne Fritzler, Neele Leithäuser

专题命中 多模态Agent :multimodal(title)

Comments 29 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19701 2025-07-29 cs.RO stat.ML 71%

PhysVarMix: Physics-Informed Variational Mixture Model for Multi-Modal Trajectory Prediction

Haichuan Li, Tomi Westerlund

机构 * Turku Intelligent Embedded and Robotics Systems lab, Faculty of Technology University of Turku(图尔库智能嵌入式与机器人系统实验室,技术学院图尔库大学)

专题命中 多模态Agent :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22034 2025-06-30 cs.RO 71%

Multi-Robot Assembly of Deformable Linear Objects Using Multi-Modal Perception

Kejia Chen, Celina Dettmering, Florian Pachler, Zhuo Liu, Yue Zhang, Tailai Cheng, Jonas Dirr, Zhenshan Bing, Alois Knoll, Rüdiger Daub

机构 * Chair of Robotics, Artificial Intelligence and Real-time Systems, School of Computation, Information and Technology(机器人学、人工智能与实时系统教授会,计算、信息与技术学院) Institute for Machine Tools and Industrial Management, School of Engineering and Design(机械加工与工业管理研究所,工程与设计学院)

专题命中 多模态Agent :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15010 2025-05-20 cs.CY cs.HC cs.LG 71%

ChatISA: A Prompt-Engineered, In-House Multi-Modal Generative AI Chatbot for Information Systems Education

Fadel M. Megahed, Ying-Ju Chen, Joshua A. Ferris, Cameron Resatar, Kaitlyn Ross, Younghwa Lee, L. Allison Jones-Farmer

专题命中 多模态Agent :multi-modal(title)

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16446 2025-04-24 cond-mat.mtrl-sci cond-mat.mes-hall cond-mat.soft 71%

Mumott -- a Python package for the analysis of multi-modal tensor tomography data

Leonard C. Nielsen, Mads Carlsen, Sici Wang, Arthur Baroni, Torne Tänzer, Marianne Liebi, Paul Erhart

专题命中 多模态Agent :multi-modal(title)

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14163 2025-02-21 cs.HC 71%

"It Brought the Model to Life": Exploring the Embodiment of Multimodal I3Ms for People who are Blind or have Low Vision

Samuel Reinders, Matthew Butler, Kim Marriott

专题命中 多模态Agent :multimodal(title)

Comments Conditionally Accepted to appear at ACM CHI Conference on Human Factors in Computing Systems (CHI '25), April 26 - May 1, 2025, Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13081 2025-01-23 cs.CR 71%

Real-Time Multi-Modal Subcomponent-Level Measurements for Trustworthy System Monitoring and Malware Detection

Farshad Khorrami, Ramesh Karri, Prashanth Krishnamurthy

专题命中 多模态Agent :multi-modal(title)

Comments 12 pages, 29 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17282 2024-12-24 cs.RO 71%

LMD-PGN: Cross-Modal Knowledge Distillation from First-Person-View Images to Third-Person-View BEV Maps for Universal Point Goal Navigation

Riku Uemura, Kanji Tanaka, Kenta Tsukahara, Daiki Iwata

专题命中 多模态Agent :cross-modal(title)

Comments Draft version of a conference paper: 5 pages with 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17751 2024-09-27 stat.ME stat.AP stat.CO 71%

Granger Causality for Mixed Time Series Generalized Linear Models: A Case Study on Multimodal Brain Connectivity

Luiza S. C. Piancastelli, Wagner Barreto-Souza, Norbert J. Fortin, Keiland W. Cooper, Hernando Ombao

专题命中 多模态Agent :multimodal(title)

Comments Paper submitted for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06222 2024-08-13 cs.HC 71%

ARCADE: An Augmented Reality Display Environment for Multimodal Interaction with Conversational Agents

Carolin Schindler, Daiki Mayumi, Yuki Matsuda, Niklas Rach, Keiichi Yasumoto, Wolfgang Minker

专题命中 多模态Agent :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19053 2024-07-30 cs.SE 71%

A Study of Using Multimodal LLMs for Non-Crash Functional Bug Detection in Android Apps

Bangyan Ju, Jin Yang, Tingting Yu, Tamerlan Abdullayev, Yuanyuan Wu, Dingbang Wang, Yu Zhao

专题命中 多模态Agent :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18374 2024-04-30 cs.RO 71%

Trajectory Optimization for Adaptive Informative Path Planning with Multimodal Sensing

Joshua Ott, Edward Balaban, Mykel Kochenderfer

专题命中 多模态Agent :multimodal(title)

Comments IEEE International Conference on Control, Decision and Information Technologies

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03088 2024-01-09 cs.RO cs.HC 71%

The RoSiD Tool: Empowering Users to Design Multimodal Signals for Human-Robot Collaboration

Nathaniel Dennler, David Delgado, Daniel Zeng, Stefanos Nikolaidis, Maja Matarić

专题命中 多模态Agent :multimodal(title)

Comments Accepted to ISER 2023. 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00097 2023-10-03 math.OC 71%

A multimodal tourist trip planner integrating road and pedestrian networks

Tommaso Adamo, Lucio Colizzi, Giovanni Dimauro, Gianpaolo Ghiani, Emanuela Guerriero

专题命中 多模态Agent :multimodal(title)

Journal ref Expert Systems with applications 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08907 2023-06-16 q-bio.BM cs.LG 71%

MCPI: Integrating Multimodal Data for Enhanced Prediction of Compound Protein Interactions

Li Zhang, Wenhao Li, Haotian Guan, Zhiquan He, Mingjun Cheng, Han Wang

专题命中 多模态Agent :multimodal(title)

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03870 2023-06-07 physics.med-ph physics.optics q-bio.TO 71%

Multimodal imaging of the mouse eye using visible light photoacoustic ophthalmoscopy and near-infrared-II optical coherence tomography

Richard Haindl, Valentina Bellemo, Praveenbalaji Rajendran, Bingyao Tan, Mengyang Liu, Qifa Zhou, Rainer A. Leitgeb, Wolfgang Drexler, Leopold Schmetterer, Manojit Pramanik

专题命中 多模态Agent :multimodal(title)

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10814 2023-05-24 cs.GT cs.RO math.OC 71%

MPOGames: Efficient Multimodal Partially Observable Dynamic Games

Oswin So, Paul Drews, Thomas Balch, Velin Dimitrov, Guy Rosman, Evangelos A. Theodorou

专题命中 多模态Agent :multimodal(title)

Comments Accepted to ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11602 2022-11-22 cs.LG cs.HC cs.MA 71%

Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback

Josh Abramson, Arun Ahuja, Federico Carnevale, Petko Georgiev, Alex Goldin, Alden Hung, Jessica Landon, Jirka Lhotka, Timothy Lillicrap, Alistair Muldal, George Powell, Adam Santoro, Guy Scully, Sanjana Srivastava, Tamara von Glehn, Greg Wayne, Nathaniel Wong, Chen Yan, Rui Zhu

专题命中 多模态Agent :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.07928 2022-03-04 cs.RO 71%

Complex In-Hand Manipulation via Compliance-Enabled Finger Gaiting and Multi-Modal Planning

Andrew S. Morgan, Kaiyu Hang, Bowen Wen, Kostas Bekris, Aaron M. Dollar

专题命中 多模态Agent :multi-modal(title)

Comments IEEE Robotics and Automation Letters, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏