LEMON: Local Explanations via Modality-aware OptimizatioN
LEMON:通过模态感知优化实现局部解释
机构 * University of Bristol(布里斯托大学)
专题命中 图文多模态 :multimodal(abstract)
AI总结 LEMON是一种高效的多模态局部解释框架,通过模态感知优化生成统一解释,减少计算成本并提升解释忠实性。
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
LEMON:通过模态感知优化实现局部解释
机构 * University of Bristol(布里斯托大学)
专题命中 图文多模态 :multimodal(abstract)
AI总结 LEMON是一种高效的多模态局部解释框架,通过模态感知优化生成统一解释,减少计算成本并提升解释忠实性。
知识向量削弱:高效无训练卸载方法用于大视觉-语言模型
机构 * Sogang University(ソガン大学) ; New York University(纽约大学)
专题命中 图文多模态 :multimodal(abstract)
AI总结 KVW提出了一种无需训练的高效卸载方法,通过削弱模型中被激活的知识向量,有效防止模型利用有害知识,提升计算效率。
接触丰富机器人任务的安全学习:从经典学习方法到安全基础模型的综述
机构 * Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genova, Italy(人类-机器人接口与交互实验室,意大利技术研究院,热那亚,意大利) ; Ph.D. program of national interest in Robotics and Intelligent Machines (DRIM) and Università di Genova, Genoa, Italy(机器人与智能机器国家利益博士项目(DRIM)和热那亚大学,热那亚,意大利) ; Edwardson School of Industrial Engineering, Purdue University, West Lafayette, IN, USA(工业工程埃德华森学校,普渡大学,西拉法伊斯,美国)
专题命中 图文多模态 :multimodal(abstract)
AI总结 本文综述了接触丰富机器人任务的安全学习方法,探讨了从经典学习方法到安全基础模型的发展,分析了安全探索与执行的关键技术及未来方向。
Comments version 2
FedSM: 用于在长尾数据联邦学习中减少偏见的语义引导特征混合
专题命中 图文多模态 :image-text(abstract)
AI总结 FedSM通过语义引导的特征混合和轻量级分类器重训练,有效减少联邦学习中长尾数据的偏见问题。
Journal ref IEEE Internet of Things Journal, 2026
RLLaVA: 一种面向语言和视觉助手的强化学习中心框架
机构 * SKLCCSE, Institute of Artificial Intelligence, Beihang University, Beijing, China(信息与电子技术学院,人工智能研究院,北京航空航天大学,北京,中国) ; Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) ; Hangzhou International Innovation Institute, Beihang University, Hangzhou, China(杭州国际创新研究院,北京航空航天大学,杭州,中国)
专题命中 图文多模态 :multi-modal(abstract)
AI总结 RLLaVA 提出了一种强化学习中心框架,通过解耦算法逻辑与模型架构,实现高效训练和多任务扩展,提升视觉-语言模型性能。
Comments The code is available at https://github.com/TinyLoopX/RLLaVA
聚焦:一种高效的视觉-语言模型流式集中架构
专题命中 图文多模态 :cross-modal(abstract)
AI总结 Focus提出了一种高效的视觉-语言模型流式集中架构,通过分层压缩和细粒度冗余消除,实现2.4倍速度提升和3.3倍能效提升。
Comments HPCA 2026
视觉语言推理中测试时扩展的极限与收益
专题命中 图文多模态 :multimodal(abstract)
AI总结 研究探讨了测试时扩展在视觉语言推理中的效果,发现其在不同任务和模型上表现不一,需根据具体需求定制策略。
Comments Mohammadjavad Ahmadpour and Amirmadhi Meighani contributed equally to this work
令牌扩展-合并:面向视觉-语言-动作模型的免训练 令牌压缩
机构 * College of Science and Engineering, Hamad Bin Khalifa University(1 科学与工程学院,哈马德·本·卡西姆大学) ; Mohamed bin Zayed University of Artificial Intelligence(2 摩萨·本·扎耶德人工智能大学) ; College of Computer Science and Technology, Zhejiang University(3 计算机科学与技术学院,浙江大学)
专题命中 图文多模态 :multimodal(abstract)
AI总结 TEAM-VLA通过动态令牌扩展与合并机制,实现无需训练的视觉-语言-动作模型高效推理,提升速度并保持任务性能。
Comments 8 pages, 5 figures
GLaD:面向视觉-语言-动作模型的几何潜在蒸馏
机构 * MBZUAI ; University of Illinois Chicago(伊利诺伊大学芝加哥分校)
专题命中 图文多模态 :multimodal(abstract)
AI总结 GLaD通过引入几何意识的预训练机制,提升了视觉-语言-动作模型的空间推理和策略泛化能力,无需依赖深度传感器或3D标注。
RobustVLA: 为视觉-语言-动作模型引入鲁棒性感知的强化学习后训练
机构 * Westlake University(西湖大学)
专题命中 图文多模态 :multi-modal(abstract)
AI总结 RobustVLA通过引入鲁棒性感知的强化学习后训练方法,提升视觉-语言-动作模型在环境不确定性下的鲁棒性和可靠性。
机构 * Tsinghua University(清华大学) ; Goertek Inc(歌尔股份有限公司)
专题命中 图文多模态 :multimodal(abstract)
Comments 23 pages, 13 figures, 7 tables
机构 * School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学) ; School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学) ; National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学) ; School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学)
专题命中 图文多模态 :multi-modal(abstract)
机构 * Zhejiang University(浙江大学)
专题命中 图文多模态 :cross-modal(abstract)
专题命中 图文多模态 :image-text(abstract)
Comments 5pages,4figures
机构 * The University of Tokyo(东京大学) ; Sony Group Corporation(索尼集团公司) ; Sony AI(索尼人工智能)
专题命中 图文多模态 :multi-modal(abstract)
机构 * Nell Hodgson Woodruff School of Nursing(Nell Hodgson Woodruff护理学院) ; School of Computer Science(计算机科学学院) ; Department of Pediatrics(儿科系) ; Department of Computer Science(计算机科学系) ; Department of Epidemiology(流行病学系) ; Department of Anesthesiology(麻醉学系) ; Department of Surgery(外科系) ; Department of Medicine(医学系) ; School of Medicine(医学院)
专题命中 图文多模态 :cross-modal(abstract)
机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) ; Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州先进研究所) ; Tsinghua University(清华大学)
专题命中 图文多模态 :image-text(abstract)
Comments 13 pages
专题命中 图文多模态 :cross-modal(abstract)
机构 * Nagoya University, Japan(名古屋大学)
专题命中 图文多模态 :multimodal(abstract)
Comments To be presented at the 1st Workshop on Intelligent Cobodied Assistance and Robotic Empowerment (iCARE). 2025 Conference on Robot Learning (CoRL)
机构 * Centre for Advanced Robotics Technology Innovation (CARTIN), School of Electrical and Electronic Engineering, Nanyang Technological University(先进机器人技术创新中心(CARTIN)、电子与电气工程学院、南洋理工大学)
专题命中 图文多模态 :multi-modal(abstract)
专题命中 图文多模态 :multimodal(abstract)
机构 * Department of Mechanical and Aerospace Engineering, Tandon School of Engineering, New York University(机械与航空航天工程系,坦顿工程学院,纽约大学) ; GenAuto.ai by General Autonomy Inc.(General Autonomy Inc. 的 GenAuto.ai)
专题命中 图文多模态 :multimodal(abstract)
专题命中 图文多模态 :multimodal(abstract)
机构 * Manning College of Information \& Computer Sciences, University of Massachusetts Amherst, Amherst, U.S.
专题命中 图文多模态 :multimodal(abstract)
机构 * Intelligent Space Robotics Laboratory(智能空间机器人实验室) ; Skolkovo Institute of Science and Technology(斯克尔科沃科学与技术研究所)
专题命中 图文多模态 :multimodal(abstract)
机构 * Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology(信息科学系,科学技术研究生学校,科学技术研究所) ; Department of Electrical and Electronic Engineering, Faculty of Engineering Science, Kansai University(电气电子工程系,工学科学大学) ; Department of Electronics, Kobe City College of Technology(电子系,神户市立技术学院) ; OMRON SINIC X Corporation(OMRON SINIC X公司)
专题命中 图文多模态 :multimodal(abstract)
Comments Published in IEEE Access, Jul 14 2025
Journal ref IEEE Access, vol. 13, pp. 125420-125441, 2025
机构 * UC Berkeley(伯克利大学) ; Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
专题命中 图文多模态 :multimodal(abstract)
Comments See our blog post at https://flowreinforce.github.io
专题命中 图文多模态 :multimodal(abstract)
Comments 19 pages, 6 figures
专题命中 图文多模态 :multimodal(abstract)
Comments 14 Pages
机构 * École Polytechnique Fédérale de Lausanne(联邦理工学院洛桑分校)
专题命中 图文多模态 :MLLM(abstract)