CoverNet: Multimodal Behavior Prediction using Trajectory Sets
专题命中 多模态Agent :multimodal(title,abstract)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :multimodal(title,abstract)
专题命中 多模态Agent :multi-modal(title,abstract)
Comments accepted by the 2019 IEEE Intelligent Vehicles Symposium (IV)
专题命中 多模态Agent :multimodal(title,abstract)
Comments Copyright (c) 2019 IEEE. Personal use of this material is permitted. However, permission to use this material for any other purposes must be obtained from the IEEE by sending a request to pubs-permissions@ieee.org. Accepted to IEEE Transactions on Medical Imaging
专题命中 多模态Agent :multimodal(title,abstract)
Comments Presented at AI-HRI AAAI-FSS, 2018 (arXiv:1809.06606)
专题命中 多模态Agent :multi-modal(title,abstract)
Comments Master's thesis, 80 pages, 24 figures
专题命中 多模态Agent :multimodal(title,abstract)
专题命中 多模态Agent :multimodal(title,abstract)
专题命中 多模态Agent :multimodal(title,abstract)
Comments IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2018 -- 8 pages, 5 figures
专题命中 多模态Agent :multimodal(title,abstract)
Comments Accepted as a conference paper at KDD 2018
专题命中 多模态Agent :multimodal(title,abstract)
专题命中 多模态Agent :multimodal(title,abstract)
专题命中 多模态Agent :multimodal(title,abstract)
Comments 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society
专题命中 多模态Agent :multimodal(title,abstract)
Comments Scaling Up Reinforcement Learning (SURL) Workshop @ European Machine Learning Conference (ECML)
专题命中 多模态Agent :multi-modal(title);multimodal(abstract)
Comments in Scientific Reports 2015
专题命中 多模态Agent :multimodal(title);multi-modal(abstract)
专题命中 多模态Agent :multimodal(title,abstract)
Comments Minor rewrites
专题命中 多模态Agent :multimodal(title,abstract)
Journal ref Proc. R. Soc. B (2007) 274, 347-357
专题命中 多模态Agent :multimodal(title,comments);分类 cs.CL、cs.AI
Comments propaganda, hate-speech, disinformation, misinformation, fake news, LLMs, GPT-4, multimodality, multimodal LLMs
面向基于视觉-语言-动作(VLA)的端到端自动驾驶的协同多模态交互
机构 * National University of Singapore (NUS)(新加坡国立大学) ; Hunan University(湖南大学) ; The University of Western Australia (UWA)(西澳大学)
专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文针对现有VLA自动驾驶模型决策不可靠、多模态交互不足的问题,提出含三类核心组件的多模态交互与多轨迹规划系统,实验显示其在安全推理与场景感知上优于现有系统。
CharTool: 集成工具的视觉推理用于图表理解
机构 * X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院X-LANCE实验室) ; Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室) ; Suzhou Laboratory(苏州实验室) ; AISpeech Co., Ltd.(思必驰科技股份有限公司)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI
AI总结 本文提出CharTool,通过集成工具提升多模态大语言模型对图表的理解能力,通过双重数据管道和代理强化学习,在六个图表基准测试中取得显著提升。
Comments Accepted by ACMMM 2026
ForgeryVCR: 通过高效的取证工具在MLLMs中实现视觉中心推理用于图像伪造检测与定位
机构 * Shenzhen University(深圳大学) ; Tencent Youtu Lab(腾讯优图实验室)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 ForgeryVCR通过高效的取证工具实现视觉中心推理,提升图像伪造检测与定位的性能。
无幻觉的GUI定位:基于无回归的布局感知匹配
机构 * School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI
AI总结 该研究提出无回归框架,通过解耦指令理解与布局感知定位,在ScreenSpot-Pro和Mind2Web数据集上显著提升GUI定位的准确率、成功率及元素选择率,抑制了坐标幻觉。
UI-MOPD:用于持续GUI智能体学习的多平台策略蒸馏
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Xiaomi(小米) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) ; Zhejiang University(浙江大学) ; Peng Cheng Laboratory(鹏城实验室)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 针对构建多平台GUI智能体的挑战,构建高质量数据集Uni-GUI,提出UI-MOPD方法,通过多教师策略蒸馏实现持续学习,动态选教师并转移行为先验,实验证明其能平衡跨平台能力保留与新平台适应。
Comments Technical report. 27 pages, 7 figures, 7 tables
PhysAgent:用于可靠远程心率估计的多智能体框架
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 PhysAgent是一种推理时多智能体候选验证框架,以多个基础rPPG估计器输出为待验证生理假设,结合Qwen3-VL-4B多模态大语言模型推理与确定性融合,提升远程心率估计的稳定性与可靠性。
Beacon:智能体何时及如何执行智能体视觉推理
机构 * Peking University(北京大学) ; Kling Team(Kling团队) ; HKUST(GZ)(香港科技大学(广州)) ; CUHK(香港中文大学) ; ZJU(浙江大学) ; THU(清华大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 本研究针对现有智能体视觉推理模型模式适应性有限、工具增益被损害抵消的问题,提出Beacon模型,通过强化学习相关机制提升性能与适应性,在多基准上表现优异。
Comments 33 pages
超越缩放:学习用于超高分辨率遥感的多工具视觉推理
机构 * National University of Defense Technology(国防科技大学) ; Wuhan University(武汉大学) ; Tsinghua University(清华大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 针对超高分辨率遥感图像给多模态大语言模型带来的挑战,提出GeoMTVR数据集,结合监督微调与以工具注意力为重点的强化学习算法开发GeoLens,实验证明其在多工具视觉推理方面优于直接推理和单工具放大基线。
自进化智能体图像恢复:通过深思熟虑的规划与直觉执行
机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; School of Life Sciences, Tsinghua University(清华大学生命科学学院) ; Intelligent Science & Technology Academy of CASIC(中国科学院 CASIC 智能科学与技术学院)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 提出SEAR框架,将图像恢复建模为序列决策问题,采用直觉执行器与深思熟虑规划器,结合剪枝感知蒙特卡洛树搜索和自进化情景记忆,解决贪婪搜索和信息利用不足问题。
DocArena:将原始文档转化为可控的训练环境用于文档搜索代理
机构 * Rochester Institute of Technology(罗切斯特理工学院) ; Adobe Research(Adobe研究院)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 提出DocArena自动化流程,通过多模态文档结构化、推理型QA对构建和质量控制,生成可控训练环境,使基于文本LLM的搜索代理在多模态文档检索和问答中取得最佳性能。
Comments search agent for documents
OmniSpace: 自动驾驶多模态大语言模型的高效几何感知
机构 * University of Arkansas(阿肯色大学) ; Google Research, Google(谷歌研究院) ; University of Liverpool(利物浦大学) ; Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究所)
专题命中 多模态Agent :MLLM(summary_cn);multimodal(abstract);分类 cs.CV
AI总结 提出OmniSpace,一种即插即用的几何感知范式,通过相机位姿注入器、多视图极线注意力模块和3D几何蒸馏目标,从纯2D观测中提升MLLM的空间推理能力,在多个自动驾驶基准上超越现有方法。
HDRAgent: 一种用于多曝光HDR成像的智能体框架
机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) ; Shenzhen Research Institute, Northwestern Polytechnical University(西北工业大学深圳研究院) ; Zhejiang University(浙江大学) ; Camera Group, DJI(大疆相机部门)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 提出首个智能体驱动的HDR成像框架HDRAgent,通过细粒度上下文知识匹配、感知-失真反馈机制和智能体引导的生成对齐策略,自适应选择重建策略,减少复杂动态场景中的鬼影和局部伪影。