arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-04 至 2026-03-04 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 13 篇

2603.02888 2026-03-04 cs.CV 79%

LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval

LLandMark:一种用于地标感知多模态交互视频检索的多智能体框架

Minh-Chi Phung, Thien-Bao Le, Cam-Tu Tran-Thi, Thu-Dieu Nguyen-Thi, Vu-Hung Dao

机构 * AI VIETNAM(AI越南)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 LLandMark通过多智能体框架实现地标感知的多模态视频检索,结合LLM和OCR技术提升文化场景下的检索能力与可解释性。

Comments Accepted by AAAI 2026 Workshop on New Frontiers in Information Retrieval

Journal ref AAAI 2026 Workshop on New Frontiers in Information Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02626 2026-03-04 cs.AI 79%

See and Remember: A Multimodal Agent for Web Traversal

见与忆:一种用于网页浏览的多模态代理

Xinjun Wang, Shengyao Wang, Aimin Zhou, Hao Hao

机构 * Shanghai Institute of AI for Education(上海人工智能教育研究院) East China Normal University(华东师范大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 V-GEMS通过视觉 grounding 和显式记忆系统实现精确稳健的网页浏览,实验显示其在导航任务中性能提升28.7%

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02519 2026-03-04 cs.MM cs.IR 79%

Agentic Mixed-Source Multi-Modal Misinformation Detection with Adaptive Test-Time Scaling

具有自适应测试时间缩放的代理混合源多模态虚假信息检测

Wei Jiang, Tong Chen, Wei Yuan, Quoc Viet Hung Nguyen, Hongzhi Yin

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.MM

AI总结 AgentM3D通过自适应测试时间缩放和多代理框架提升零样本多模态虚假信息检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20346 2026-03-04 cs.CR cs.AI cs.LG 79%

Multimodal Multi-Agent Ransomware Analysis Using AutoGen

基于AutoGen的多模态多智能体勒索软件分析

Asifullah Khan, Aimen Wadood, Mubashar Iqbal, Umme Zahoora

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出基于AutoGen的多模态多智能体框架,通过融合静态、动态和网络信息,提升勒索软件分类的准确性和稳定性。

Comments 46 pages, 11 figures and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02635 2026-03-04 cs.LG 78%

SaFeR-ToolKit: Structured Reasoning via Virtual Tool Calling for Multimodal Safety

SaFeR-ToolKit: 通过虚拟工具调用实现多模态安全的结构化推理

Zixuan Xu, Tiancheng He, Huahui Yi, Kun Wang, Xi Chen, Gongli Xi, Qiankun Li, Kang Li, Yang Liu, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Beijing University of Posts and Telecommunications(北京邮电大学) West China Hospital, Sichuan University(四川大学华西医院) Nanyang Technological University(南洋理工大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 SaFeR-ToolKit通过虚拟工具调用实现多模态安全的结构化推理,提升安全性、帮助性和推理严谨性,同时保持通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02503 2026-03-04 eess.SY cs.SY 78%

Joint Estimation of Dynamic O-D Demand and Choice Models for Dynamic Multi-modal Networks: Computational Graph-Based Learning and Hypothesis Tests

动态多模式网络中动态O-D需求与选择模型的联合估计:基于计算图的学习与假设检验

Xiaoyu Ma, Sean Qian

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本研究提出基于计算图的学习方法,联合估计多模式网络中动态O-D需求与选择模型,通过假设检验框架提升模型的统计显著性分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02366 2026-03-04 cs.HC cs.AI 74%

PlayWrite: A Multimodal System for AI Supported Narrative Co-Authoring Through Play in XR

PlayWrite: 一种通过XR中的游戏化互动实现AI支持的叙事协作写作的多模态系统

Esen K. Tütüncü, Qian Zhou, Frederik Brudy, George Fitzmaurice, Fraser Anderson

机构 * Institute of Neurosciences of the University of Barcelona(巴塞罗那大学神经科学研究所) Autodesk Research(Autodesk研究)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 PlayWrite是一种通过XR中的游戏化互动实现AI支持的叙事协作写作的多模态系统,通过直接操控虚拟角色和道具,结合多智能体AI管道和大型语言模型,促进高度即兴和游戏化的创作过程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03198 2026-03-04 cs.RO cs.CL cs.CV 73%

ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments

ACE-Brain-0:空间智能作为通用具身化体系的共享框架

Ziyang Gong, Zehang Luo, Anke Tang, Zhe Liu, Shi Fu, Zhi Hou, Ganlin Yang, Weiyun Wang, Xiaofeng Wang, Jianbo Liu, Gen Luo, Haolan Kang, Shuang Luo, Yue Zhou, Yong Luo, Li Shen, Xiaosong Jia, Yao Mu, Xue Yang, Chunxiao Liu, Junchi Yan, Hengshuang Zhao, Dacheng Tao, Xiaogang Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) University of Science(科学技术大学) Fudan University(复旦大学) Xiamen University(厦门大学) East China Normal University(华东师范大学) Wuhan University(武汉大学) Sun Yat-sen University(中山大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

AI总结 ACE-Brain-0 通过空间智能作为共享框架,统一了自动驾驶、机器人和 UAVs 的具身化任务,采用 SSR 范式和 GRPO 方法实现跨领域泛化和领域精通的平衡。

Comments Code: https://github.com/ACE-BRAIN-Team/ACE-Brain-0 Hugging Face: https://huggingface.co/ACE-Brain/ACE-Brain-0-8B

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23141 2026-03-04 cs.CV 57%

Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents

Earth-Agent: 解锁地球观测的全貌与潜力

Peilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang, Xinjie Huo, Jinhua Yu, Wanghan Xu, Wenlong Zhang, Lei Bai, Conghui He, Weijia Li

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

AI总结 Earth-Agent是一种结合RGB和光谱数据的代理框架,通过多模态工具生态系统实现跨模态、多步骤推理,提升地球观测分析的科学性和应用潜力。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17520 2026-03-04 cs.RO cs.CV 57%

InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation

InstructVLA: 从理解到操作的视觉-语言-动作指令微调

Shuai Yang, Hao Li, Bin Wang, Yilun Chen, Yang Tian, Tai Wang, Hanqing Wang, Feng Zhao, Yiyi Liao, Jiangmiao Pang

机构 * University of Science and Technology of China(中国科学技术大学) Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 InstructVLA通过视觉-语言-动作指令微调,实现了在多模态推理和动作生成之间的平衡,提升了机器人在复杂任务中的操控性能和泛化能力。

Comments 48 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03121 2026-03-04 cs.SE 50%

RippleGUItester: Change-Aware Exploratory Testing

RippleGUItester: 基于变更的探索性测试

Yanqi Su, Michael Pradel, Chunyang Chen

专题命中 多模态Agent :multimodal(abstract)

AI总结 RippleGUItester通过基于大语言模型的变更影响分析,结合多模态bug检测,发现由代码变更引入的bug,提升测试效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15602 2026-03-04 cs.NE 50%

Estimate Hitting Time by Hitting Probability for Elitist Evolutionary Algorithms

通过击中概率估计精英进化算法的击中时间

Jun He, Siang Yew Chong, Xin Yao

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出通过击中概率的漂移分析方法,用于估计精英进化算法的击中时间,通过引入路径处理多模态适应度景观,简化了击中概率的计算并比较了两种算法的性能。

Journal ref IEEE Transactions on Evolutionary Computation 12 November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00079 2026-03-04 cs.HC 50%

Enhancing the Interpretability of SHAP Values Using Large Language Models

利用大语言模型增强SHAP值的可解释性

Xianlong Zeng, Kewen Zhu

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文利用大语言模型提升SHAP值的可解释性,使非技术用户更易理解模型预测。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏