IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments
IndoorUAV:在连续室内环境中进行视觉-语言无人机导航的基准测试
专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.AI
AI总结 IndoorUAV通过构建室内无人机视觉-语言导航基准,结合多模态推理和任务分解,为长视距和短视距导航提供高质量数据与模型支持。
视觉与机器人
视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。
IndoorUAV:在连续室内环境中进行视觉-语言无人机导航的基准测试
专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.AI
AI总结 IndoorUAV通过构建室内无人机视觉-语言导航基准,结合多模态推理和任务分解,为长视距和短视距导航提供高质量数据与模型支持。
多LLM协作用于药物推荐
机构 * Computer Science Laboratory, SRI International(SRI国际计算机科学实验室) ; University of Maryland St. Joseph Medical Center(马里兰大学圣约瑟夫医疗中心)
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
AI总结 本文提出基于LLM化学的多模型协作方法,通过增强互补性、稳定性和校准性,提高药物推荐的可靠性与可信度。
Comments 8 pages, 5 figures, 1 table
SPARK: 面向模拟的关节化重建与VLK知识
专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.CV
AI总结 SPARK通过结合VLK和生成扩散模型,实现从单张图像中生成物理一致的关节化3D物体,提升机器人操作和交互建模的应用效果。
Comments Project page: https://heyumeng.com/SPARK/index.html. 17 pages, 7 figures
Saga:从大量未标记IMU数据中捕捉多粒度语义以实现用户感知
机构 * Shanghai Jiao Tong University(上海交通大学) ; Donghua University(东华大学)
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
AI总结 Saga通过预训练和贝叶斯优化实现从大量未标记IMU数据中捕捉多粒度语义,仅用少量标记数据即可达到高用户感知准确率。
Comments 2025 IEEE 45th International Conference on Distributed Computing Systems (ICDCS)
Journal ref Proceedings of the IEEE International Conference on Distributed Computing Systems (ICDCS), 2025
MonoDream:基于全景梦境的单目视觉-语言导航
机构 * Horizon Robotics
专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.CV
AI总结 MonoDream通过轻量级VLA框架和潜在全景梦境任务,提升单目视觉-语言导航的性能,缩小与全景方法的差距。
机构 * Department of Mechanical Engineering, Institute of Science Tokyo(科学东京研究院机械工程系) ; Satellite Research and Development, Interstellar Technologies Inc.(星际技术公司卫星研发部) ; Department of Advanced Energy, The University of Tokyo(东京大学先进能源系) ; Department of Spacecraft Engineering, Japan Aerospace Exploration Agency(日本宇宙航空研究开发机构航天器工程系)
专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.LG
Comments Submitted to IEEE Robotics and Automation Letters
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
Comments 31 Pages, 16 Figures, 9 Tables
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Lab(上海人工智能实验室) ; Nanyang Technological University(南洋理工大学)
专题命中 VLA模型 :vision language action(abstract);分类 cs.RO、cs.CV
机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息处理国家重点实验室,计算机学院,北京大学) ; Hong Kong University of Science and Technology(香港科技大学) ; National University of Singapore(新加坡国家大学) ; Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.CV
机构 * School of Mechanical Engineering, Beijing Institute of Technology(机械工程学院,北京理工大学)
专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI
专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI
机构 * College of Transportation, Tongji University(同济大学交通运输学院)
专题命中 VLA模型 :VLA(abstract);分类 cs.RO、cs.CV
专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI
机构 * Dept. of Informatics, University of Hamburg(信息学院,汉堡大学)
专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI
Comments Accepted and published at the 34th International Conference on Artificial Neural Networks (ICANN 2025)
机构 * Zhejiang University(浙江大学) ; Westlake University(西湖大学) ; UCAS-Terminus AI Lab(UCAS-terminus人工智能实验室) ; Victoria University of Wellington(威灵顿维多利亚大学)
专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.LG
机构 * University of Glasgow(格拉斯哥大学) ; University of Leeds(利兹大学) ; University of Warwick(沃里克大学)
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
专题命中 VLA模型 :vision-language-action(abstract);分类 cs.RO、cs.AI
Comments We withdraw our submission following peer review feedback that identified methodological limitations: specifically, our experimental design does not adequately support the causal claims made in the submission. The work was preliminary undergraduate research that requires substantial additional experimental validation to properly establish the proposed causal relationships
机构 * Uppsala University(乌普萨拉大学) ; Florida Institute of Technology(佛罗里达理工学院)
专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI
Comments 6 pages, 2 figures
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
机构 * School of Computer Science, Peking University(北京大学计算机科学学院) ; Center for Data Science, Peking University(北京大学数据科学中心) ; National Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学通用人工智能国家重点实验室) ; Northwestern University(西北大学) ; Yale University(耶鲁大学)
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
Comments Accepted at ICML 2025
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
Comments Accepted to ICML 2025
机构 * Center for Secure & Intelligent Critical Systems, Old Dominion University, Virginia, USA(安全与智能关键系统中心,旧 Dominion 大学,弗吉尼亚州,美国) ; School of Cybersecurity, Old Dominion University, Virginia, USA(网络安全学院,旧 Dominion 大学,弗吉尼亚州,美国)
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
Comments Accepted at IEEE IWCMC. 6 pages, 4 Figures, 3 tables
机构 * Czech Institute of Informatics, Robotics and Cybernetics(捷克信息学、机器人学与自动控制研究所) ; Czech Technical University in Prague(布拉格捷克技术大学)
专题命中 VLA模型 :vision-language-action(abstract);分类 cs.RO、cs.LG
Comments 7 pages, 5 figures, 2 tables, conference
Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
机构 * Department of Civil and Environmental Engineering, University of Wisconsin-Madison(土木与环境工程系,威斯康星大学麦迪逊分校) ; Lyles School of Civil and Construction Engineering, Purdue University(莱尔斯土木与建设工程学院,普渡大学) ; Elmore Family School of Electrical and Computer Engineering, Purdue University(埃尔摩家族电气与计算机工程学院,普渡大学) ; Department of Computer Science and Engineering, The University of Texas at Arlington(计算机科学与工程系,德克萨斯大学阿灵顿分校) ; Google(谷歌)
专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI
Comments 14 pages, 7 figures
专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.AI
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
专题命中 VLA模型 :action model(abstract);分类 cs.CV、cs.LG
专题命中 VLA模型 :action model(abstract);分类 cs.AI、cs.LG
Comments 34 pages, 3 figures, 2 tables