MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation
MetaWorld-X: 通过VLM协调的专家实现人形机器人的分层世界建模
Yutong Shen, Hangxu Liu, Penghui Liu, Jiashuo Luo, Yongkang Zhang, Rex Morvley, Chen Jiang, Jianwei Zhang, Lei Zhang
机构
*
University of Hamburg(汉堡大学)
;
School of Information Science and Technology, Beijing University of Technology(信息科学与技术学院,北京理工大学)
;
School of Information Science and Engineering, Fudan University(信息科学与工程学院,复旦大学)
;
University of Alberta(阿尔伯塔大学)
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
DeAR: 通过分解注意力头角色实现细粒度VLM适应
Yiming Ma, Hongkun Yang, Lionel Z. Wang, Bin Chen, Weizhi Xian, Jianzhi Teng
机构
*
Chongqing Research Institute of Harbin Institute of Technology(哈尔滨工业大学重庆研究所)
;
Ocean University of China(中国海洋大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong, China(计算机科学与工程系,香港中文大学,香港,中国)
;
Institute of Medical Intelligence and XR, The Chinese University of Hong Kong, Hong Kong, China(医学智能与XR研究所,香港中文大学,香港,中国)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.CV
SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
SPEX:一种用于光谱遥感图像土地覆盖提取的视觉-语言模型
Dongchen Si, Di Wang, Erzhong Gao, Xiaolei Qin, Liu Zhao, Jing Zhang, Minqiang Xu, Jianbo Zhan, Jianshe Wang, Lin Liu, Bo Du, Liangpei Zhang
机构
*
College of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)
;
iFlytek Co., Ltd.(iFlytek公司)
;
National Engineering Research Center of Speech and Language Information Processing(语音与语言信息处理国家工程研究中心)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Zhongguancun Academy(中关村学院)
;
National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心)
;
Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(湖北省多媒体与网络通信工程重点实验室,武汉大学)
;
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(测绘遥感信息工程国家重点实验室,武汉大学)
机构
*
Tongji University(同济大学)
;
University of California, Santa Cruz(加州大学圣克ruz分校)
;
Amazon(亚马逊)
;
East China Normal University(华东师范大学)
;
Shanghai Eye Disease Prevention and Treatment Center(上海眼病预防与治疗中心)
机构
*
School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院)
;
School of Computer Science and Technology, Tiangong University(天津工大学计算机科学与技术学院)
;
Key Research Center for Surface Monitoring and Analysis of Relics, State Administration of Cultural Heritage(文物表面监测与分析关键研究中心,国家文物局)
机构
*
Tongji University(同济大学)
;
The City University of New York(纽约城市大学)
;
University of Technology Sydney(悉尼大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
Shenzhen University of Advanced Technology(深圳先进技术大学)
专题命中
VLM训练与架构
:grounding(abstract);multimodal large language model(abstract);分类 cs.AI
CDE: Concept-Driven Exploration for Reinforcement Learning
CDE:基于概念的强化学习探索
Le Mao, Andrew H. Liu, Renos Zabounidis, Yanan Niu, Zachary Kingston, Joseph Campbell
机构
*
Department of Electrical and Computer Engineering, Purdue University(电子工程系,普渡大学)
;
Department of Computer Science, Purdue University(计算机科学系,普渡大学)
;
Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)
;
Department of Management of Technology, EPFL(技术管理系,瑞士联邦理工学院)
ActivePose: Active 6D Object Pose Estimation and Tracking for Robotic Manipulation
ActivePose: 用于机器人操作的主动6D物体姿态估计与跟踪
Sheng Liu, Zhe Li, Weiheng Wang, Han Sun, Heng Zhang, Hongpeng Chen, Yusen Qin, Arash Ajoudani, Yizhao Wang
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Istituto Italiano di Tecnologia(意大利理工学院)
;
The Hong Kong Polytechnic University(香港理工大学)
;
D-Robotics(D-机器人)
Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding
具有4D空间-时间嵌入的自监督多模态世界模型
Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati, Patrick Alan Johnson, Favyen Bastani, Jeff Rose, James Ryan Hennessy, Robert Guralnick, Douglas Soltis, Pamela Soltis, Shaowen Wang
机构
*
Ecological Intelligence Lab(生态智能实验室)
;
School of Complex Adaptive Systems(复杂适应系统学院)
;
University of Houston(休斯顿大学)
;
Geosensing Systems Engineering & Sciences Lab(传感系统工程与科学实验室)
;
Stanford University(斯坦福大学)
;
Allen Institute for Artificial Intelligence(人工智能研究院)
;
Spatial Intelligence Lab(空间智能实验室)
;
Department of Computer Science(计算机科学系)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Florida Museum of Natural History(佛罗里达自然历史博物馆)
;
University of Florida(佛罗里达大学)
;
NSF Institute for Geospatial Understanding(国家科学基金会地理理解研究所)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
隐形微调:通过自动生成的CoT打破RVLMs的对齐
Le Yu, Zhengyue Zhao, Yawen Zheng, Yunhao Liu
机构
*
Machine Intelligence Laboratory, Sichuan University(四川大学人工智能实验室)
;
University of Wisconsin--Madison(威斯康星大学麦迪逊分校)
;
Department of Automation, Tsinghua University(清华大学自动化系)
;
Global Innovation Exchange, Tsinghua University(清华大学全球创新交流中心)