V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments 12 pages, 5 figures, 3 tables
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments 12 pages, 5 figures, 3 tables
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL
Comments ICML 2024
基于反思引导的在线策略自蒸馏的测试时自演进GUI视觉定位
专题命中 多模态Agent :MLLM(summary_cn,abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 针对现有GUI视觉定位模型部署后难适配未见界面且无法反思失败探索的问题,提出测试时自演进框架,结合MLLM反思器、反思引导的在线策略自蒸馏等方法,在六个基准上平均提升7.4%准确率,完善了GUI智能体自演进能力。
ReMMD: 面向多模态虚假信息检测的现实多语言多图像智能体验证
机构 * Shanghai Jiaotong University(上海交通大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Tsinghua University(清华大学) ; Central South University(中南大学) ; China Electronics Technology Group Corporation 15th Research Institute(中国电子科技集团公司第十五研究所)
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
AI总结 提出ReMMD框架,包含多语言多图像基准ReMMDBench和持久记忆验证器ReMMD-Agent,通过原子点分解和可重用证据集实现高效准确的多模态虚假信息检测。
Comments The project is available at https://dang-ai.github.io/ReMMD
POINTS-Seeker:从零开始训练多模态代理搜索模型
机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) ; CMIC, Shanghai Jiao Tong University(上海交通大学CMIC) ; WeChat AI, Tencent(腾讯WeChat AI)
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文提出POINTS-Seeker模型,通过引入代理播种机制和V-Fold压缩方案,解决长周期交互中的证据检索问题,并在六个基准测试中超越现有模型。
Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索
机构 * School of Artificial Intelligence UCAS(中国科学院大学人工智能学院) ; Institute of Automation CAS(中国科学院自动化研究所) ; Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团) ; RUC(中国人民大学) ; BIT(北京理工大学)
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
AI总结 提出Visual-Seeker,一种通过主动视觉推理进行视觉原生多模态深度搜索的智能体,在五个基准上达到最先进性能,甚至超越专有模型。
M$^3$Exam: 面向真实用户-智能体交互的多模态记忆基准
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; Beijing University of Chemical Technology(北京化工大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) ; Beijing Institute of Technology (Zhuhai)(北京理工大学(珠海)) ; Tencent Hy(腾讯(深圳)) ; Peng Cheng Laboratory(鹏城实验室)
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
AI总结 提出M$^3$Exam基准,用于评估多模态大语言模型在真实用户-智能体交互中的跨模态推理和隐式信息推断能力,并设计M$^3$Proctor方法通过按需处理视觉源提升准确率13%,同时降低索引构建时间和检索token超70%。
Demo2Tutorial:从人类经验到多模态软件教程
机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)
专题命中 多模态Agent :multimodal(title,abstract);image-text(abstract);分类 cs.CV
AI总结 提出Demo2Tutorial框架,通过屏幕录制和交互日志将人类经验解析为结构化多模态教程,用于人类学习和GUI智能体训练,实验证明其生成质量超越人工教程并提升任务效率。
Comments Accepted by CVPR 2026
MementoGUI: 学习代理多模态记忆控制以实现长周期GUI代理
机构 * University of Rochester(罗切斯特大学) ; MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 多模态Agent :multimodal(title);MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 本文提出MementoGUI,一种学习代理多模态记忆控制框架,用于提升长周期GUI代理的任务状态维持能力,通过模块化记忆控制和可扩展的数据管道提高记忆检索和决策效率。
Comments Preprint, 15 pages, 4 figures, 5 tables
基于触觉的多模态融合在具身智能中的应用:视觉、语言和接触驱动范式的综述
机构 * School of Electronic Science and Engineering, Xi’an Jiaotong University, China(西安交通大学电子科学与技术学院) ; Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州)人工智能研究所) ; State Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室) ; Purple Mountain Laboratory, China(紫金山实验室) ; Institute for Math & AI, Wuhan University, China(武汉大学数学与人工智能学院) ; Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University, Australia(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院) ; School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China(北京邮电大学人工智能学院) ; Institute of Big Data, Fudan University, China(复旦大学大数据研究院) ; Linkerbot (Beijing) Technology Co., Ltd, China(北京链动科技有限公司) ; School of Engineering, Swinburne University of Technology, Melbourne(斯威本技术大学工程学院)
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文综述了多模态触觉融合在具身智能中的研究,探讨了如何通过整合视觉、语言和触觉信息来提升物理交互与语义推理的结合,提出了一种分层的分类体系,并总结了当前的研究挑战和未来方向。
Comments 20 pages, 8 figures
跨模态导航与多智能体强化学习
机构 * Khoury College of Computer Sciences(计算机科学学院)
专题命中 多模态Agent :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI
AI总结 本文提出CRONA框架,通过多智能体强化学习实现跨模态导航,利用辅助信念和集中式多模态批评者提升协作效率,实验表明多智能体方法在视觉-听觉导航中优于单智能体基线。
重新思考多模态问答中的信息合成:多智能体视角
机构 * Arizona State University(亚利桑那州立大学)
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
AI总结 本文提出多智能体框架MAMMQA,通过分解查询、跨模态推理和整合答案,提升多模态问答的准确性和可解释性。
代理到代理网络中的模态本原路由:一种多模态A2A协议扩展
机构 * AI Architect, Author—Data Engineering for Multimodal AI (O’Reilly)(人工智能架构师,作者—多模态AI的数据工程(O’Reilly)) ; Stanford School of Engineering (April 2026)(斯坦福大学工程学院(2026年4月))
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
AI总结 本文提出MMA2A架构,通过多模态本原路由提升任务准确率,其在跨模态任务中表现优于文本瓶颈基线,尤其在视觉依赖任务中效果显著,但增加了1.8倍的延迟。
Comments 14 pages, 4 figures (TikZ). PDFLaTeX. Supplementary code and experiment artifacts: https://github.com/vasundras/modality-native-routing-a2a-protocol
MONETA:通过地理信息和多智能体系统进行多模态行业分类
机构 * Trustworthy Human Language Technologies(可信人类语言技术实验室) ; Technical University of Darmstadt, Germany(德国达姆施塔特工业大学) ; Deutsche Bundesbank(德国联邦银行) ; Frankfurt University of Applied Sciences, Germany(德国法兰克福应用技术大学) ; Research Center for Trustworthy Data Science and Security, Ruhr University Bochum, Germany(德国波鸿鲁尔大学可信数据科学与安全研究中心)
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 本文提出MONETA,首个基于文本和地理空间数据的多模态行业分类基准,利用多智能体系统提升分类精度,实现62.10%和74.10%的分类性能。
Comments Accepted to ACL 2026 Main Conference
Commander-GPT:多模态讽刺检测的分工与路由
专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI
AI总结 本文提出Commander-GPT框架,通过分工协作的LLM代理团队,结合三种指挥官类型,提升多模态讽刺检测性能,实验结果显示在F1分数上优于现有方法。
MultiPress:一种用于可解释多模态新闻分类的多智能体框架
机构 * New York Institute of Technology(纽约理工学院) ; University of Arizona(亚利桑那大学) ; University of Macau(澳门大学) ; Peking University(北京大学) ; Juniata College(朱尼亚塔学院)
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
AI总结 本文提出MultiPress多智能体框架,通过多模块协作和检索增强推理提升多模态新闻分类的准确性和可解释性。
Comments Accepted in International Joint Conference on Neural Networks (IJCNN) 2026
ATP-Bench: 向多模态大语言模型交错生成的代理工具规划迈进
机构 * Qwen Large Model Application Team, Alibaba(阿里巴巴通义千问大模型应用团队) ; Huazhong University of Science and Technology(华中科技大学) ; Zhejiang University(浙江大学)
专题命中 多模态Agent :MLLM(title,abstract);multimodal(abstract);分类 cs.AI
AI总结 针对多模态大语言模型交错生成中事实性与创造性难以统一的问题,提出ATP-Bench基准,包含7702个问题-答案对,评估代理工具规划能力,揭示模型在交错规划中的不足。
Webscraper:利用多模态大语言模型进行索引-内容网页抓取
机构 * Dept. of Information Management, National Taiwan University, Taipei, Taiwan(国立台湾大学资讯管理学系,台北,台湾)
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 本文提出Webscraper框架,利用多模态大语言模型自动导航交互界面并提取结构化数据,通过五阶段提示和定制工具提升动态网站抓取准确性。
FusionAgent:一种具有动态模型选择的多模态代理用于人体识别
机构 * Michigan State University(密歇根州立大学) ; University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 FusionAgent通过动态模型选择提升人体识别鲁棒性,利用多模态大语言模型进行样本特定的模型选择,结合ACT分数融合方法,实验表明其在效率和性能上优于现有方法。
Comments CVPR 2026
GeoSketch: 一种基于神经符号的方法,用于几何多模态推理中的辅助线构造与仿射变换
机构 * Fudan University(复旦大学) ; IFLYTEK CO.LTD(若lytek有限公司) ; Zhejiang University(浙江大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; National University of Defense Technology(国防科技大学) ; Hainan University(海南大学)
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 GeoSketch通过整合感知、符号推理和绘图动作模块,实现动态几何推理,提升多模态推理的准确性和解决问题的成功率。
VistaWise: 构建低成本代理的跨模态知识图谱用于Minecraft
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; University of Queensland(昆士兰大学) ; Tencent(腾讯)
专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI
AI总结 VistaWise通过整合跨模态知识图谱和专用模型,实现低成本、高效率的Minecraft代理构建。
Comments Accepted by EMNLP 2025 main
SiMO: 单模态可操作的多模态协作感知
机构 * Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统上海研究院) ; School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) ; School of Mechatronic Engineering and Automation, Shanghai University(上海大学机械电子工程与自动化学院)
专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV
AI总结 SiMO通过单模态可操作的多模态协作感知方法,解决多模态特征融合中的语义不匹配问题,提升协作感知性能。
Comments Accepted to ICLR 2026. This arXiv version includes an additional appendix (Appendix 15) containing further philosophical discussion not included in the official ICLR peer-reviewed version
VTool-R1: 通过在多模态工具使用上的强化学习使VLMs学会通过图像思考
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Michigan Ann Arbor(密歇根大学安娜堡分校) ; Independent Researcher(独立研究者)
专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI
AI总结 VTool-R1通过强化学习训练视觉语言模型生成多模态思考链,提升其通过图像进行推理的能力。
Comments ICLR 2026
Chain-of-Caption: 无需训练提升多模态大语言模型的指称表达理解
机构 * Queen Mary University of London(伦敦女王学院)
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出无需训练的Chain-of-Caption框架,通过结合多种上下文提升多模态大语言模型在指称表达理解任务中的性能。
Comments 4 pages, 5 figures, 2 tables
事实还是假象?评估深度伪造检测器在多模态虚假信息检测中的作用
机构 * The University of Western Australia(西澳大学) ; Fatih Sultan Mehmet Vakif University(法提赫·苏丹·梅赫梅特·瓦基夫大学) ; Aljazeera Media Network Investigative Department(半岛电视台调查部门)
专题命中 多模态Agent :multimodal(title,abstract);image-text(abstract);分类 cs.CV
AI总结 研究评估了深度伪造检测器在多模态虚假信息检测中的作用,发现其独立价值有限,而证据驱动的事实核查系统表现更优。
COSINT-Agent: 一种面向中文开源情报的知识驱动多模态智能体
专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 COSINT-Agent通过整合多模态大语言模型与实体-事件-场景知识图谱,提升中文开源情报的多模态推理能力与情报生成效率。
Comments This manuscript (arXiv:2503.03215) is being withdrawn at the supervisor's request. The content is preliminary and needs further internal revision and approval before public release. We will resubmit a revised version after completion. Apologies for the inconvenience
DriveMLM: 通过行为规划状态对齐多模态大语言模型以实现自动驾驶
机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) ; Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心)
专题命中 多模态Agent :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV
AI总结 DriveMLM通过多模态大语言模型对齐行为规划状态,提升自动驾驶系统的决策能力与闭环性能。
Comments Accepted to Visual Intelligence
Journal ref Visual Intelligence, Volume 3, article number 22, (2025)
DrivePI: 基于空间感知的4D MLLM用于统一自动驾驶理解、感知、预测与规划
机构 * The University of Hong Kong(香港大学) ; Yinwang Intelligent Technology Co. Ltd.(英维智能科技有限公司) ; Tianjin University(天津大学) ; Huazhong University of Science and Technology(华中科技大学)
专题命中 多模态Agent :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV
AI总结 DrivePI是一种基于空间感知的4D MLLM,用于统一自动驾驶的理解、感知、预测和规划,通过端到端优化实现多任务并行,提升性能并减少碰撞率。
AgenticCyber: 一种基于生成式AI的多智能体系统,用于多模态威胁检测与自适应响应在网络安全中
机构 * TNTech
专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
AI总结 AgenticCyber通过生成式AI驱动的多智能体系统,实现多模态威胁检测与自适应响应,提升网络安全性能和态势感知能力。
Comments 6 pages for IEEE conference
InEx:通过内省与跨模态多智能体协作缓解幻觉
专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV
AI总结 InEx通过内省推理和跨模态多智能体协作,自主缓解大型语言模型的幻觉问题,实验表明其在多个基准上表现优异。
Comments Published in AAAI 2026