arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2797 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2797 篇

2602.21394 2026-04-23 cs.CR 78%

MemoPhishAgent: Memory-Augmented Multi-Modal LLM Agent for Phishing URL Detection

MemoPhishAgent: 基于记忆的多模态大语言模型代理用于钓鱼URL检测

Xuan Chen, Hao Liu, Tao Yuan, Mehran Kafai, Piotr Habas, Xiangyu Zhang

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出MemoPhishAgent,通过整合多模态推理与记忆机制,提升钓鱼URL检测性能,实验显示其在召回率上优于现有方法,并在真实场景中实现高召回率。

Comments ACL 2026 Industry Track Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19081 2026-04-22 cs.SE 78%

Proactive Detection of GUI Defects in Multi-Window Scenarios via Multimodal Reasoning

通过多模态推理主动检测多窗口场景中的GUI缺陷

Xinyao Zhang, Rui Wang, Jinhao Cui, Haotian Huang, Wei Xue, Wenhua Hu, Jianwen Xiang, Rui Hao

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出一种端到端框架,通过主动触发多窗口状态,利用多模态大语言模型检测定位GUI缺陷,实验表明多窗口设置显著增加布局缺陷暴露,方法在应用和细粒度层面均优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17459 2026-04-21 cs.IR 78%

Transparent and Controllable Recommendation Filtering via Multimodal Multi-Agent Collaboration

通过多模态多智能体协作实现透明且可控的推荐过滤

Chi Zhang, Zhipeng Xu, Jiahao Liu, Dongsheng Li, Hansu Gu, Peng Zhang, Ning Gu, Tun Lu

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出一种结合端到云协作、多模态感知和多智能体编排的框架,有效减少推荐系统中的过度关联和误报,提升用户控制和透明度。

Comments 14 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03509 2026-04-20 cs.RO 78%

Sampling-Based Multi-Modal Multi-Robot Multi-Goal Path Planning

基于采样的多模态多机器人多目标路径规划

Valentin N. Hartmann, Tirza Heinle, Yijiang Huang, Stelian Coros

机构 * Computational Robotics Lab ETH Zürich(机器人计算实验室ETH Zurich)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出基于采样的多机器人多目标路径规划方法,解决多机器人协同任务中的路径规划问题,实现概率完备和渐进最优的规划算法。

Comments 25 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10501 2026-04-14 cs.CR 78%

MuSimA: A Tool with Multi-modal Input for Generating Bespoke ABAC Datasets

MuSimA:一种支持多模输入的生成定制ABAC数据集的工具

Saket Jha, Karthikeya S. M. Yelisetty, Singabattu Sathya, Shamik Sural

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出MuSimA工具,支持多模输入生成定制ABAC数据集,利用大语言模型自动提取分布参数,用于测试ABAC系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06914 2026-04-09 cs.LG 78%

Equivariant Multi-agent Reinforcement Learning for Multimodal Vehicle-to-Infrastructure Systems

等价多智能体强化学习用于多模态车-基础设施系统

Charbel Bou Chaaya, Mehdi Bennis

机构 * Centre for Wireless Communications, University of Oulu(奥卢大学无线通信中心)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文研究了车-基础设施系统,通过分布式基站收集多模态数据,采用去中心化速率最大化方法,利用等价性对称性提出自监督学习框架,提升多模态感知和强化学习效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18836 2026-04-08 cs.RO 78%

Multimodal Classification Network Guided Trajectory Planning for Four-Wheel Independent Steering Autonomous Parking Considering Obstacle Attributes

四轮独立转向自主泊车的多模态分类网络引导轨迹规划考虑障碍属性

Jingjia Teng, Yang Li, Yougang Bian, Manjiang Hu, Yingbai Hu, Guofa Li, Jianqiang Wang

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出基于多模态感知网络和混合A*搜索的轨迹规划框架,通过考虑障碍属性提升四轮独立转向车辆在狭窄环境中的导航效率和安全性。

Comments The manuscript in this current form requires substantial revision. For this reason, I request the withdrawal of the submission to allow for comprehensive improvement before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29409 2026-04-01 cs.RO 78%

CLaD: Planning with Grounded Foresight via Cross-Modal Latent Dynamics

CLaD:通过跨模态潜在动力学实现具有基础前瞻性的规划

Andrew Jeong, Jaemin Kim, Sebin Lee, Sung-Eui Yoon

机构 * KAIST(韩国科学技术院)

专题命中 多模态Agent :cross-modal(title,abstract)

AI总结 CLaD通过不对称跨注意力机制联合建模本体和语义状态演化,利用自监督目标和EMA编码器防止表征崩溃,实现高成功率的长序列动作生成。

Comments Project page: https://andrewwwj.github.io/clad

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27202 2026-03-31 cs.PL cs.DC 78%

Sal: Multi-modal Verification of Replicated Data Types

Sal:多模态验证复制数据类型

Pranav Ramesh, Vimala Soundarapandian, KC Sivaramakrishnan

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 Sal通过结合内核可验证的自动化、SMT辅助自动化及AI辅助交互式定理证明,实现CRDT和MRDT的多模态验证,69%的验证条件由内核验证自动化解决,自动生成反例暴露隐秘错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21418 2026-03-31 cs.MA 78%

FUAS-Agents: Autonomous Multi-Modal LLM Agents for Treatment Planning in Focused Ultrasound Ablation Surgery

FUAS-代理:自主多模态大语言模型代理用于聚焦超声消融手术的治疗计划

Lina Zhao, Zihao Bian, Qingyue Chen, Yafang Li, Zhiyi Luo, Jiaxing Bai, Guangbo Li, Min He, Kezhi Li, Huaiyuan Yao, Zongjiu Zhang

专题命中 多模态Agent :multi-modal(title);multimodal(abstract)

AI总结 本文提出FUAS-代理,利用大语言模型的多模态理解和工具使用能力,通过整合患者资料和MRI数据,生成个性化治疗计划,提升临床决策效率和可靠性。

Comments 35 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25195 2026-03-27 cs.HC 78%

On-Demand Instructional Material Providing Agent Based on MLLM for Tutoring Support

基于MLLM的按需教学材料提供代理用于辅导支持

Takumi Kato, Masato Kikuchi, Tadachika Ozono

专题命中 多模态Agent :MLLM(title);multimodal(abstract)

AI总结 本文提出基于多模态大语言模型的代理,用于在一对一辅导中按需提供教学材料,通过分析对话自动检索相关图片,实验显示检索时间减少44.4秒,85.7%的试次提供可接受质量的图片。

Comments The 20th International Conference on E-Service and Knowledge Management (ESKM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22507 2026-03-24 cs.NI cs.MA eess.SP 78%

A Unified Cloud-Edge-Terminal Framework for Multimodal Integrated Sensing and Communication

多模态感知与通信一体化的统一云-边-终端框架

Yubo Peng, Luping Xiang, Kun Yang, Feibo Jiang, Kezhi Wang, Christos Masouros

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出统一云-边-终端框架,通过多模态感知与通信融合,解决异构融合、通信开销和系统扩展性等挑战,提升任务导向的多模态感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01013 2026-03-24 cs.LG 78%

TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop

TimeXL:基于LLM的多模态时间序列预测可解释方法

Yushan Jiang, Wenchao Yu, Geon Lee, Dongjin Song, Kijung Shin, Wei Cheng, Yanchi Liu, Haifeng Chen

机构 * School of Computing, University of Connecticut(大学计算机学院) Data Science & System Security Department, NEC Labs America(数据科学与系统安全部,NEC美国实验室) Kim Jaechul Graduate School of AI, KAIST(金 Jaechul人工智能研究生院,韩国科学技术院)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 TimeXL通过集成原型时间序列编码器与三个协作LLM,提升时间序列预测的准确性与可解释性,实验证明在四个真实数据集上AUC提升达8.9%。

Comments NeurIPS 2025 camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14540 2026-03-17 cs.RO 78%

Multimodal Belief-Space Covariance Steering with Active Probing and Influence for Interactive Driving

多模态信念空间协方差操控与主动探测及影响的交互驾驶

Devodita Chakravarty, John Dolan, Yiwei Lyu

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出一种多模态信念空间协方差操控方法,通过主动探测和影响策略,在交互驾驶中提升安全性和决策效率。

Comments Accepted to IEEE International Conference on Robotics and Automation (ICRA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22961 2026-03-13 stat.ME 78%

Measuring capacities in multimodal maritime port systems with anchorage queues

多模式港口系统中锚泊队列容量的测量

Debojjal Bagchi, Kyle Bathgate, Kenneth N. Mitchell, Magdalena I. Asborno, Marin M. Kress, Stephen D. Boyles

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出了一种方法,用于估算多模式港口系统的运营和终极容量,通过队列模型和微分方程模型分析休斯顿港的吞吐量及瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21100 2026-03-12 astro-ph.GA astro-ph.IM 78%

Disk Wind Feedback from High-mass Protostars. V. Application of Multi-Modal Machine Learning to Characterize Outflow Properties

高质恒星喷流反馈。V. 多模式机器学习在喷流性质表征中的应用

Duo Xu, Ioana A. Stelea, Joshua S. Speagle, Yichen Zhang, Jonathan C. Tan

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出多模式深度学习方法,通过结合空谱信息表征喷流性质,克服投影偏差,提升高质恒星形成研究的可解释性与鲁棒性。

Comments ApJ accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08336 2026-03-10 cs.RO 78%

Hierarchical Multi-Modal Planning for Fixed-Altitude Sparse Target Search and Sampling

分层多模态规划用于固定高度稀疏目标搜索与采样

Lingpeng Chen, Yuchen Zheng, Apple Pui-Yi Chui, Junfeng Wu, Ziyang Hong

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 HIMoS通过分层多模态规划提高稀疏目标搜索与采样任务的效率,结合全局与局部规划器优化路径,平衡多种传感任务。

Comments 8 pages, 9 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04363 2026-03-05 cs.RO 78%

ManipulationNet: An Infrastructure for Benchmarking Real-World Robot Manipulation with Physical Skill Challenges and Embodied Multimodal Reasoning

ManipulationNet: 一个用于基于现实世界机器人操作的基准测试基础设施,包含物理技能挑战和具身多模态推理

Yiting Chen, Kenneth Kimble, Edward H. Adelson, Tamim Asfour, Podshara Chanrungmaneekul, Sachin Chitta, Yash Chitambar, Ziyang Chen, Ken Goldberg, Danica Kragic, Hui Li, Xiang Li, Yunzhu Li, Aaron Prather, Nancy Pollard, Maximo A. Roa-Garzon, Robert Seney, Shuo Sha, Shihefeng Wang, Yu Xiang, Kaifeng Zhang, Yuke Zhu, Kaiyu Hang

机构 * Rice University(里士大学) U.S. National Institute of Standards and Technology(美国国家标准与技术研究院) Massachusetts Institute of Technology(麻省理工学院) Karlsruhe Institute of Technology(卡尔斯鲁厄技术大学) Autodesk Research(Autodesk研究) University of California, Berkeley(加州大学伯克利分校) KTH Royal Institute of Technology(皇家理工学院) Tsinghua University(清华大学) Columbia University(哥伦比亚大学) ASTM International(美国材料与试验协会) Carnegie Mellon University(卡内基梅隆大学) German Aerospace Center (DLR)(德国航空航天中心) University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Texas at Austin(德克萨斯大学奥斯汀分校) NVIDIA Research(NVIDIA研究)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 ManipulationNet通过标准化硬件和统一软件客户端,为机器人操作提供现实世界基准测试,促进物理技能和具身推理能力的系统性发展。

Comments 32 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02635 2026-03-04 cs.LG 78%

SaFeR-ToolKit: Structured Reasoning via Virtual Tool Calling for Multimodal Safety

SaFeR-ToolKit: 通过虚拟工具调用实现多模态安全的结构化推理

Zixuan Xu, Tiancheng He, Huahui Yi, Kun Wang, Xi Chen, Gongli Xi, Qiankun Li, Kang Li, Yang Liu, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Beijing University of Posts and Telecommunications(北京邮电大学) West China Hospital, Sichuan University(四川大学华西医院) Nanyang Technological University(南洋理工大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 SaFeR-ToolKit通过虚拟工具调用实现多模态安全的结构化推理,提升安全性、帮助性和推理严谨性,同时保持通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02503 2026-03-04 eess.SY cs.SY 78%

Joint Estimation of Dynamic O-D Demand and Choice Models for Dynamic Multi-modal Networks: Computational Graph-Based Learning and Hypothesis Tests

动态多模式网络中动态O-D需求与选择模型的联合估计:基于计算图的学习与假设检验

Xiaoyu Ma, Sean Qian

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本研究提出基于计算图的学习方法,联合估计多模式网络中动态O-D需求与选择模型,通过假设检验框架提升模型的统计显著性分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21157 2026-03-02 cs.RO 78%

HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning

HALO:一种用于具身多模态推理的统一视觉-语言-动作模型

Quanxin Shou, Fangqi Zhu, Shawn Chen, Puxin Yan, Zhengyang Yan, Yikun Miao, Xiaoyi Pang, Zicong Hong, Ruikai Shi, Hao Huang, Jie Zhang, Song Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 HALO提出了一种统一的视觉-语言-动作模型,通过结合文本推理、视觉子目标预测和增强的动作预测,实现具身多模态链式推理,提升了机器人在复杂环境中的表现和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19346 2026-02-24 cs.RO cs.SY eess.SY 78%

Design and Control of Modular Magnetic Millirobots for Multimodal Locomotion and Shape Reconfiguration

模块化磁性微机器人多模态运动与形状重构设计与控制

Erik Garcia Oyono, Jialin Lin, Dandan Zhang

机构 * Department of Bioengineering, Imperial College London(帝国理工学院生物工程系)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本研究提出一种模块化磁性微机器人平台,通过多模块协同实现多模态运动与形状重构,展示了在受限环境中稳健控制的潜力。

Comments Accepted by 2026 ICRA

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08417 2026-02-17 cs.RO 78%

ORACLE-Grasp: Zero-Shot Affordance-Aligned Robotic Grasping using Large Multimodal Models

ORACLE-Grasp: 基于大多模态模型的零样本 affordance 对齐机器人抓取

Avihai Giuili, Rotem Atari, Avishai Sintov

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 ORACLE-Grasp 利用大多模态模型实现零样本抓取,通过语义对齐提升抓取准确性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11057 2026-02-12 cs.LG 78%

Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language Models

分割、调和、然后征服:利用多模态语言模型解决多商品流问题

Xinyu Yuan, Yan Qiao, Zonghui Wang, Wenzhi Chen

机构 * Zhejiang University(浙江大学) Hefei University of Technology(合肥工业大学) Co-corresponding authors(共同通讯作者)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 Pram利用多模态语言模型解决多商品流问题,通过分解和调和子问题实现高效优化,性能接近线性规划求解器且运行时间显著降低。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05671 2026-02-06 cs.HC 78%

(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences

行动中的视觉:比较远程视觉援助与多模态语音代理在检查序列中的表现

Damien Rudaz, Barbara Nino Carreras, Sara Merlino, Brian L. Due, Barry Brown

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 研究比较了远程视觉援助与多模态语音代理在检查任务中的表现,发现代理无法产生基于环境的视觉动作,从而缺乏关键资源。

Comments Conditionally accepted at CHI 2026, 32 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04157 2026-02-05 cs.RO 78%

A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling

一种面向情境具身人机对话的现代系统配方,结合实时多模态大语言模型与工具调用

Dong Won Lee, Sarah Gillet, Louis-Philippe Morency, Cynthia Breazeal, Hae Won Park

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出了一种结合实时多模态大语言模型与工具调用的系统配方,用于提升情境具身人机对话的交互质量与效率。

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19185 2026-02-03 math.OC 78%

Distributionally Robust Optimization with Multimodal Decision-Dependent Ambiguity Sets

具有多模决策依赖不确定集的分布鲁棒优化

Xian Yu, Beste Basciftci

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出了一种基于ϕ-分歧度的多模决策依赖分布鲁棒优化框架,通过引入多模不确定集和分解算法,改进了DRO模型的求解方法,并通过案例验证了多模性对优化性能的提升作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03777 2026-01-08 stat.ME cs.SY eess.SY 78%

Multi-agent Optimization of Non-cooperative Multimodal Mobility Systems

非合作多模式移动系统的多智能体优化

Md Nafees Fuad Rafi, Zhaomiao Guo

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出了一种多智能体优化框架,用于分析非合作多模式移动系统中旅行者和司机的市场互动,通过均衡定价平衡供需并优化系统效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00696 2026-01-05 cs.LG cs.GT cs.RO 78%

Bayesian Inverse Games with High-Dimensional Multi-Modal Observations

高维多模态观测下的贝叶斯逆游戏

Yash Jain, Xinjie Liu, Lasse Peters, David Fridovich-Keil, Ufuk Topcu

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Delft University of Technology(代尔夫特理工大学) Sunrise Setting Ltd SAGE Publications Ltd(SAGE出版社有限公司)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract)

AI总结 本文提出了一种基于贝叶斯推断的逆博弈框架,利用多模态观测数据实时生成隐藏智能体目标的后验分布,提升推断质量并实现更安全的决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06115 2026-01-05 cs.RO cs.SY eess.SY 78%

Hybrid A* Path Planning with Multi-Modal Motion Extension for Four-Wheel Steering Mobile Robots

四轮转向移动机器人多模态运动扩展的混合A*路径规划

Runjiao Bao, Lin Zhang, Tianwei Niu, Haoyu Yuan, Shoukun Wang

机构 * School of Automation, Beijing Institute of Technology, Beijing 100081, China(自动化学院,北京理工大学,北京100081,中国)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出了一种针对四轮转向移动机器人的混合A*路径规划方法,通过多模态运动扩展提升复杂环境下的路径规划性能。

Comments Updated method details, parameters, and experimental scenarios

详情

展开后加载摘要…

URL PDF HTML 收藏