Evolving from Tool User to Creator via Training-Free Experience Reuse in Multimodal Reasoning
从工具使用者到创作者的进化:通过无训练经验重用在多模态推理中
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 UCT框架通过无训练经验重用,使智能体从工具使用者转变为工具创造者,提升多模态推理能力。
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
从工具使用者到创作者的进化:通过无训练经验重用在多模态推理中
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 UCT框架通过无训练经验重用,使智能体从工具使用者转变为工具创造者,提升多模态推理能力。
面向多模态材料数据和工作流的本体对齐结构化与重用,以实现自动重现
专题命中 多模态Agent :multimodal(title);分类 cs.AI
AI总结 本文提出一种基于本体驱动和大型语言模型的框架,用于自动提取和结构化多模态材料数据和工作流,以提高计算结果的可重现性和重用性。
Comments 39 pages, 7 figures
通过输入预测和误位修正加速多模态大语言模型游戏性能
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
AI总结 通过输入预测和误位修正方法,显著降低多模态大语言模型游戏性能的推理延迟,提升整体控制效果。
Comments UIUC 25 Fall CS 498
机构 * Department of Computer Science(计算机科学系)
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments Improved experimental setup
机构 * University of Oxford(牛津大学) ; King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学) ; Technical University of Munich(慕尼黑技术大学) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; Simon Fraser University(西蒙·弗雷泽大学) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
Comments 2nd version update to Jun.2025
机构 * Georgia Institute of Technology Atlanta, GA, USA(佐治亚理工学院)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; Nanyang Technological University(南洋理工大学) ; University of Chinese Academy of Sciences(中国科学院大学)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
机构 * Center for Intelligent Machines, McGill University, Montreal, Canada(麦吉尔大学智能机器中心,加拿大蒙特利尔) ; Mila - Quebec AI institute, Montreal, Canada(魁北克AI研究所)
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
Comments 9 pages, 3 figures, International Conference on Medical Image Computing and Computer-Assisted Intervention
机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) ; Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) ; Tsinghua University(清华大学) ; Huazhong University of Science and Technology(华中科技大学)
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
Comments Accepted by the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)
机构 * Stanford University(斯坦福大学) ; Microsoft Research(微软研究院)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
Comments 8 pages, 3 figures
专题命中 多模态Agent :multimodal(title);分类 cs.AI
专题命中 多模态Agent :multimodal(title);分类 cs.CL
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
Comments AAAI 2025 (Project page: https://twoongg.github.io/projects/flare/)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
Comments Accepted and presented at ECCV 2024 2nd Workshop on Vision-Centric Autonomous Driving (VCAD) on September 30, 2024. 13 pages, 5 figures
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
Comments Accepted to NeurIPS 2024
专题命中 多模态Agent :multimodal(title);分类 cs.AI
Comments 5 pages, 1 figure, 1 table, accepted in Embodied AI 2024 Workshop held in conjunction with CVPR 2024
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments The 1st place solution of End-to-end Driving at Scale at the CVPR 2024 Autonomous Grand Challenge
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments Accepted for publication at MICCAI 2024 workshop on AI for Imaging Genomics Learning (AIIG)
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments CO and DHP contributed equally to this work. JSD and ETR are corresponding authors
专题命中 多模态Agent :multimodal(title);分类 cs.AI
Comments Accepted as full paper in AAMAS 2024
专题命中 多模态Agent :multimodal(title);分类 cs.AI
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
Comments Challenge report doi.org/10.1016/j.cmpb.2023.107561
Journal ref Computer Methods and Programs in Biomedicine, Volume 236, 2023
专题命中 多模态Agent :cross-modal(title);分类 cs.CV
Comments Our project page with videos is at https://RchalYang.github.io/LocoTransformer
专题命中 多模态Agent :multimodal(title);分类 cs.CL
Comments 10 pages; This position paper was presented at the Rethinking the Senses: A Workshop on Multisensory Embodied Experiences and Disability Interactions associated with the ACM CHI Conference on Human Factors in Computing Systems, May 2021
Journal ref ACM CHI Conference on Human Factors in Computing Systems, May 2021
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Journal ref Proceedings of 21st International Radar Symposium (IRS 2020)
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments The paper has been accepted by IEEE Transactions on Intelligent Transportation Systems 2020
专题命中 多模态Agent :multimodal(title);分类 cs.AI
Comments 16 pages, 13 figures, new experiments, new explanatory figures for intuition and new title
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Journal ref Asian Conference on Computer Vision (ACCV). 2018. 511-526