An automatic deep learning-based workflow for glioblastoma survival prediction using pre-operative multimodal MR images
专题命中 多模态Agent :multimodal(title)
Journal ref Advances in Radiation Oncology, 2021
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :multimodal(title)
Journal ref Advances in Radiation Oncology, 2021
专题命中 多模态Agent :multi-modal(title)
Comments 3 pages, 6 figures
Journal ref Proceedings of the 6th International workshop on Image Processing for Art Investigation (IP4AI) 2018
专题命中 多模态Agent :multimodal(title)
Comments 33 pages, 8 figures
Journal ref Acta Biomaterialia 101, 459-468 (2019)
专题命中 多模态Agent :multi-modal(title)
Comments 18 pages, 31 figures
专题命中 多模态Agent :multimodal(title)
Comments 17 pages, International Journal of Robotics Research
专题命中 多模态Agent :multimodal(title)
Comments Accepted in the proceedings of 6th IEEE International Smart Cities Conference, Sep 28 - Oct 1, 2020
专题命中 多模态Agent :multi-modal(title)
Comments Ali-akbar Agha-mohammadi is the Principal Investigator. arXiv admin note: substantial text overlap with arXiv:2002.00515
专题命中 多模态Agent :multi-modal(title)
专题命中 多模态Agent :multimodal(title)
Comments 8 pages 6 figures main article; 7 pages 9 figures supporting informations
Journal ref Physical Review B 96, 035440 (2018)
专题命中 多模态Agent :multi-modal(title)
专题命中 多模态Agent :multi-modal(title)
专题命中 多模态Agent :multimodal(title)
Comments 9 pages, 7 figures
专题命中 多模态Agent :multimodal(title)
Comments 7 pages, 8 figures, changed to match version accepted by MNRAS
Journal ref Mon.Not.Roy.Astron.Soc.378:1365-1370,2007
无行为的信念:测量视觉-语言模型中心智理论到协同社会行动的转化
机构 * CIAMS, Université Paris-Saclay(巴黎萨克雷大学 CIAMS) ; CQSB, Sorbonne Université(索邦大学 CQSB)
专题命中 多模态Agent :multimodal(abstract,abstract_cn);分类 cs.AI
AI总结 该研究提出基准MOSAIC评估视觉-语言模型的心智理论到协同社会行动的转化,发现多数模型存在瓶颈,而带显式ToM模块的PCM-LLM表现优异,证实显式信念-行动耦合的作用。
Act2Intention:通过从GUI操作推断用户意图开发主动移动智能体的基准
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract_cn);分类 cs.AI
AI总结 本文提出Act2Intention框架及基准数据集,构建主动移动智能体,经实验验证该基准可显著提升智能体意图理解、预测与执行性能,为主动智能体研究提供标准化平台。
ActFER: 通过主动工具增强的视觉推理实现代理面部表情识别
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV
AI总结 ActFER通过主动视觉证据获取和多模态推理提升面部表情识别,采用UC-GRPO算法优化局部检查和情绪感知,实验显示其在AU预测准确率上显著优于传统方法。
Comments 10 pages, 7 figures
VOS-Agent:第8届LSVOS挑战赛(MOSEv2赛道)的第一名解决方案
机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) ; Nanyang Technological University(南洋理工大学) ; Shenzhen Loop Area Institute(深圳河套学院)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 针对复杂视频目标分割中微小目标、语义主导目标的稳健传播难题,本文提出VOS-Agent协作框架,以SAM3为共享密集分割模块,根据目标特征激活专用智能体,在ECCV 2026第8届LSVOS挑战赛MOSEv2赛道获第一名,J&JF指标达69.82%。
Comments 1st Place Solution for the 8th LSVOS MOSEv2 Challenge (ECCV 2026 Workshop)
SkillLens:用于检索增强型GUI动作预测与在线策略蒸馏的可视化技能卡片
专题命中 多模态Agent :multimodal(abstract,abstract_cn);分类 cs.AI
AI总结 SkillLens提出可视化技能卡片(VSCs),结合轨迹转卡片方法与检索增强机制,在两个多模态网页基准上提升了冻结GPT-5.4-mini执行器及Qwen3-VL-2B学生模型的GUI动作预测性能。
用于具身操纵的数据金字塔
机构 * PKU(北京大学) ; NTU(南洋理工大学) ; HKUST(香港科技大学) ; NUS(新加坡国立大学) ; CUHK(香港中文大学) ; HKU(香港大学) ; Duke(杜克大学) ; UCB(加州大学伯克利分校) ; GBU(未提及具体中文名的机构) ; NJU(南京大学) ; SJTU(上海交通大学)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
AI总结 研究围绕具身操纵数据生态系统展开,构建跨越五个互补数据源的“数据金字塔”,通过数据配方分析具身基础模型,将数据组成与多种能力联系起来,并讨论了六个开放挑战,为下一代具身系统设计奠定基础。
Comments Awesome Embodied Data Pyramid; Project Page at https://jasper-aaa.github.io/embodied-data-pyramid/ GitHub Repo at https://github.com/worldbench/awesome-embodied-data-pyramid
NormAct:具身规划中隐藏社会规范遵守的基准
机构 * State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI)(通用人工智能国家重点实验室,北京通用人工智能研究院) ; China Academy of Information and Communications Technology(中国信息通信研究院) ; ShanghaiTech University(上海科技大学)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI
AI总结 提出NormAct基准,评估多模态大语言模型在具身规划中遵守隐藏社会规范的能力,发现模型在67.3%情况下达成显式目标,但仅26.4%遵守隐藏规范;提出NormPerceptor方法将任务成功率从24.2%提升至46.7%。
Comments This version revises the paper content and adds substantial details on the benchmark design and experimental setup
视觉工具使用的错觉:对图像思考的因果审查
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 多模态Agent :multimodal(abstract,abstract_cn);分类 cs.AI
AI总结 本文通过因果审查方法,发现多模态大语言模型的视觉工具使用存在“调用而不查看”“查看而不规划”等错觉,整体准确率增益下大量场景无因果效果。
Metis:记忆基础模型
机构 * MemTensor (Shanghai) Technology Co., Ltd.(墨芯(上海)科技有限公司) ; Renmin University of China(中国人民大学) ; National University of Singapore(新加坡国立大学) ; Shanghai Jiao Tong University(上海交通大学) ; Tongji University(同济大学)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL
AI总结 该研究提出首个记忆基础模型原型Metis,赋予基础模型原生记忆能力,通过新架构与优化目标实现,经实验验证其具备原生记忆能力并发布相关资源。
Comments 46 pages, 11 figures, 16 tables
SeekBrain:用于加速神经科学发现的自主多智能体系统
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.AI
AI总结 SeekBrain是用于加速神经科学发现的自主多智能体框架,通过分层规划与跨模态分析生成分析流程,在BrainArena基准上优于现有智能体,实际研究中成功揭示了斑马鱼与小鼠的神经表征规律。
医学中的智能体人工智能:临床转化的架构、应用、评估及挑战
机构 * School of Software Engineering, Dalian University(大连大学软件工程学院) ; Affiliated Zhongshan Hospital of Dalian University(大连大学附属中山医院) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; University of California, San Francisco(加州大学旧金山分校) ; Yale University(耶鲁大学) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
AI总结 研究探讨医学中智能体人工智能的架构、应用等,通过范围审查筛选相关研究,指出其范围未定论且评估与临床需求不符,证据基础有限,临床转化需更清晰定义、可重复评估等。
Comments Review article, 6 figures, 2 tables. 42 pages
不应学习的地方:基于子集归因约束的先验对齐训练以实现可靠的决策制定
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) ; Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系) ; Communication University of China(中国传媒大学) ; Imperial College London(伦敦帝国学院) ; School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区网络科学与技术学院)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 本文提出了一种基于归因的先验对齐方法,通过子集选择归因技术约束模型依赖于人类先验区域,从而提升决策的可靠性。
基于工具增强证据的多智能体协作推理用于城市区域剖析
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; University of Washington(华盛顿大学) ; Chang’an University(长安大学) ; University of Wisconsin - Madison(威斯康星大学麦迪逊分校)
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.AI
AI总结 研究针对城市区域剖析问题,提出UrbanAgent框架,通过多智能体协作推理解决跨模态不一致,将指标预测扩展为闭环过程,经实验验证其性能优于现有基线,在未见城市设置中有强泛化性。
Comments Accepted by KDD 2026
一次前向胜过两次:InnerZoom 实现准确高效的 GUI 定位
机构 * Alibaba Group(阿里巴巴集团) ; Tongyi Lab(通义实验室) ; University of Technology Sydney(悉尼大学) ; Adelaide University(阿德莱德大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 提出 InnerZoom 框架,通过跨层证据桥接在单次前向传播中实现精确 GUI 元素定位,无需额外裁剪重跑,在六个基准上达到最优性能,同时降低延迟和计算量。
USS: 面向具身视觉跟踪的统一空间-语义提示与潜在动力学学习
机构 * Nanyang Technological University(南洋理工大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 提出统一空间-语义提示范式替代纯文本目标指示,设计端到端框架USS支持多种提示类型,通过潜在世界模型提升时序鲁棒性,在真实机器人实验中优于纯文本方法。
Colon-Bench:一种用于全流程结肠镜视频中可扩展密集病变标注的智能工作流
机构 * King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹科学与技术大学)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV
AI总结 提出Colon-Bench,通过多阶段智能工作流实现全流程结肠镜视频的密集病变标注,包含528个视频、14类病变、30万+边界框等,并评估多模态大模型在病变分类、开放词汇视频目标分割和视频视觉问答上的性能。
Comments published at MICCAI 2026
Skill-3D:面向智能体3D空间推理的场景感知技能进化
机构 * Zhejiang University(浙江大学) ; University of Technology Sydney(技术悉尼大学) ; OPPO Research Institute(OPPO研究院)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 提出Skill-3D框架,通过场景记忆和技能库的协同进化,使智能体根据场景自适应选择工具,显著提升3D空间推理中工具使用的正确性和充分性。