AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Project at https://github.com/OpenBMB/AppCopilot
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Project at https://github.com/OpenBMB/AppCopilot
专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Appeared in EMNLP 2025 main conference. To better understand prompt injection attacks, see https://people.duke.edu/~zg70/code/PromptInjection.pdf
Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments The paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract)
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 24 pages, 11 figures
专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
基于多智能体闭环推理的多模态光谱有机结构解析
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 本文提出多智能体系统MACROS,经海量光谱数据训练后可实现零样本泛化,提升结构解析速度与准确性,为全自动结构解析及自主实验室发展奠定基础。
截图还是工具?在混合GUI-MCP计算机使用智能体中引出工具使用并管理多模态上下文
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 该研究针对混合GUI-MCP计算机使用智能体,发现工具使用存在采用缺口,通过调整工具奖励与上下文压缩优化,提升了智能体性能并降低输入成本。
ClinLens:面向纵向多模态临床数据科学的长周期编码智能体
机构 * Shandong University(山东大学)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 该研究推出CLINLENS基准,包含200项基于5种MIMIC资源的临床可执行任务,实验显示现有编码智能体与生物医学系统在正确临床分析上存在显著差距。
ETPDesigner:用于交互式多模态电子剧院节目的多智能体编排
机构 * Shanghai University(上海大学) ; Shanghai Engineering Research Center of Motion Picture Special Effects(上海电影特效工程技术研究中心) ; Institute for Math and AI, Wuhan School of Artificial Intelligence, Wuhan University(武汉大学数学与人工智能研究院武汉人工智能学院)
专题命中 多模态Agent :multimodal(title);分类 cs.CV
AI总结 研究针对电子剧院节目设计难题,提出ETPDesigner多智能体框架,能从戏剧脚本合成高质量ETP,通过全局风格锚定机制保证一致性,还实现交互功能,经ETP-Pro基准测试验证了方法在多方面的优越性。
视觉 inception:通过多模态记忆污染在代理推荐系统中妥协长期规划
机构 * City University of Hong Kong(香港城市大学)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 本文提出视觉 inception 攻击,通过污染用户上传的图片在代理推荐系统中影响长期规划,提出 CognitiveGuard 防御框架以降低攻击风险。
Comments 17 pages, 6 figures, 16 tables
Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20846-20862, 2026
通过多模态大语言模型对策略进行视觉检查实现开放式多智能体自动课程
机构 * Sapienza University of Rome(罗马第一大学) ; Sony AI(索尼人工智能)
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
AI总结 研究强化学习中开放式多智能体自动课程设计难题,提出通过策略视觉检查(VIP)利用视频语言模型处理视频并提供课程建议,经星际争霸多智能体挑战赛实证,显示该方法能生成更有效课程。
LLM引导的规划:多模态核监管文档的多跳推理
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 提出LLM引导的规划方法,通过动态知识图谱状态和文档树工具,在多跳推理任务中实现81.5%准确率,显著优于无状态规划方法。
Comments Accepted at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond @ ICML 2026. 8 pages (main), 3 figures, 1 algorithm
面向城市电磁场地图预测的多模态条件高分辨率Transformer
机构 * Soongsil University(崇实大学) ; Cho Chun Shik Graduate School of Mobility, Korea Advanced Institute of Science and Technology(韩国科学技术院赵春植移动研究生院) ; Department of Intelligent Semiconductors, Soongsil University(崇实大学智能半导体系)
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
AI总结 提出多条件密集预测框架,利用高分辨率Transformer骨干网络,结合特征线性调制和交叉注意力机制,从建筑布局和天线配置生成500×500电磁场地图,通过复合损失函数和测试时增强提升性能。
基于无人机多模态视觉系统的边坡变形自动监测与灾害检测
机构 * South China Normal University(华南师范大学)
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
AI总结 提出一种基于无人机LiDAR的自动化边坡灾害检测流程,包括数据采集、地面点提取、单次观测灾害筛查和多次观测变形监测,实现厘米级形变量化与危险区域识别。
Comments 29 pages, 14 figures
SciOrch: 学习编排专家大语言模型以解决前沿多模态科学推理任务
机构 * Imperial College London(伦敦帝国学院) ; The Chinese University of Hong Kong(香港中文大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; University of Oxford(牛津大学) ; Shenzhen Loop Area Institute(深圳河套学院)
专题命中 多模态Agent :multimodal(title);分类 cs.CL
AI总结 提出SciOrch框架,训练轻量级8B模型编排多个前沿大语言模型,通过MCTS和GRPO优化,在科学推理任务上超越最强单模型和多智能体基线。
BotDirector:跨对称现实的多模态交互机器人讲故事
机构 * State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China(国家一般人工智能重点实验室,BIGAI,北京,中国) ; Peking University, Beijing, China(北京大学,北京,中国)
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
AI总结 提出一个结合具身交互和自然语言交互的机器人讲故事系统,利用LLM代理将儿童创建的叙事转化为自导航群体机器人的运动序列,支持灵活场景和日常物品。
Journal ref 2026 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW)
AgentGrounder:使用多模态语言模型的零样本3D视觉点云定位
机构 * Biomechatronics and Energy-Efficient Robotics (BE2R) Lab, ITMO University(生物机械与高效能机器人实验室,ITMO大学)
专题命中 多模态Agent :multimodal(title);分类 cs.CV
AI总结 提出AgentGrounder框架,通过两阶段设计(离线构建对象查找表和在线工具驱动代理)实现零样本3D视觉定位,在ScanRefer和Nr3D上分别提升2.5%和6.3%的准确率。
Comments Code: https://github.com/be2rlab/AgentGrounder
面向LLM代理出站的应用层多模态隐通道参考监控器应用
机构 * Enclawed, LLC(Enclawed公司)
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
AI总结 本文提出了一种应用层多模态隐通道参考监控器,用于检测和防止LLM代理在消息中泄露数据,通过多阶段文本管道、媒体加密器和残余容量测量来实现对隐通道的监控和管理。
一种多模态深度感知方法用于具身参照理解
机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) ; KIT Campus Transfer GmbH (KCT)(KIT校园转移有限公司) ; Istanbul Technical University(伊斯坦布尔技术大学) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 多模态Agent :multimodal(title);分类 cs.CV
AI总结 本文提出一种多模态深度感知方法,结合语言模型数据增强、深度图模态和深度感知决策模块,提升复杂环境中的参照物识别准确性。
Comments Accepted by ICASSP 2026
CRC-SAM:基于SAM的多模态结直肠癌在CT、结肠镜和组织学图像中的分割与量化
机构 * Independent researcher(独立研究者)
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
AI总结 本文提出CRC-SAM框架,实现跨结肠镜、CT和病理图像的结直肠癌分割,通过LoRA层实现高效领域迁移,实验显示在多个数据集上优于现有方法。
Comments 4 pages, 3 figures, ISBI 2026 oral presentation
AutoGUI-v2:一个全面的多模态GUI功能理解基准
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; New Laboratory of Pattern Recognition(模式识别新实验室) ; State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) ; Hong Kong Institute of Science & Innovation(香港科学创新研究院) ; PolyU ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
AI总结 AutoGUI-v2通过多平台截图递归解析生成多样化任务,评估深度GUI功能理解和交互预测能力,揭示VLMs在功能接地与描述上的差异及复杂交互逻辑的挑战。
Comments Technical Report
Eco-Bee:一种面向校园生态系统的个性化多模态代理,用于提升学生气候意识和可持续行为
机构 * University of Hull(赫尔大学) ; The Spaceship Academy(太空学院) ; City St George’s, University of London(伦敦大学城市学院) ; King’s College London(伦敦国王学院) ; Lancaster University(兰卡斯特大学) ; University of Oxford(牛津大学)
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
AI总结 Eco-Bee通过整合大语言模型、行星边界框架(Eco-Score)和对话代理,为学生提供个性化反馈和行为激励,推动校园可持续发展。
PosterGen:基于多智能体LLM的美观化论文到海报生成
机构 * Stony Brook University(石溪大学) ; New York University(纽约大学) ; University of British Columbia(不列颠哥伦比亚大学) ; Zhejiang University(浙江大学) ; University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
AI总结 PosterGen通过多智能体LLM实现论文到海报的美观化生成,包含四个协作专业代理,提升设计质量和视觉吸引力。
Comments Project Website: https://Y-Research-SBU.github.io/PosterGen
CodeDance: 一种动态集成工具的多模态大语言模型用于可执行视觉推理
机构 * Beihang University(北京航空航天大学) ; Westlake University(西湖大学) ; ByteDance Singapore(字节跳动新加坡) ; ByteDance China(字节跳动中国)
专题命中 多模态Agent :MLLM(title);分类 cs.CV
AI总结 CodeDance通过可执行代码实现视觉推理,结合多工具协作与自检机制,提升复杂任务的灵活性和可解释性,实验显示其在多个基准测试中优于现有方法。
Comments CVPR 2026. Project page: https://codedance-vl.github.io/
基于浏览器的开源多模态内容验证助手
机构 * University of Sheffield(谢菲尔德大学) ; AFP Medialab(法新社媒体实验室)
专题命中 多模态Agent :multimodal(title);分类 cs.CL
AI总结 本文提出VERIFICATION ASSISTANT,一种浏览器工具,整合多个NLP服务,为非专家用户提供清晰的可信度信号和反虚假信息指导。
心灵超越空间:多模态大语言模型能否进行心理导航?
机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) ; Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(清华大学-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学计算机科学与技术系) ; School of Automation Science and Electrical Engineering , Beihang University(北京航空航天大学自动化科学与电气工程学院) ; college of AI, Tsinghua University(清华大学人工智能学院) ; Department of Computer Science and Technology , Tsinghua University(清华大学计算机科学与技术系)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 本文提出Video2Mental基准测试,评估多模态大语言模型的空间导航能力,发现标准预训练模型无法自然生成空间表示,NavMind通过显式认知地图提升导航性能。
PlayWrite: 一种通过XR中的游戏化互动实现AI支持的叙事协作写作的多模态系统
机构 * Institute of Neurosciences of the University of Barcelona(巴塞罗那大学神经科学研究所) ; Autodesk Research(Autodesk研究)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 PlayWrite是一种通过XR中的游戏化互动实现AI支持的叙事协作写作的多模态系统,通过直接操控虚拟角色和道具,结合多智能体AI管道和大型语言模型,促进高度即兴和游戏化的创作过程。
在简单视觉规划任务中多模态大语言模型推理的分布外泛化
机构 * Tübingen AI Center -- University of Tübingen(图宾根人工智能中心 -- 图宾根大学) ; EPFL(瑞士联邦理工学院) ; ELLIS Institute Finland -- Aalto University(芬兰ELLIS研究所 -- 阿尔托大学)
专题命中 多模态Agent :multimodal(title);分类 cs.CV
AI总结 研究多模态大语言模型在简单视觉规划任务中推理能力的分布外泛化,发现结合多种文本格式的推理方法效果最佳,纯文本模型表现优于图像输入模型。