Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
机构 * Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Northwestern Polytechnical University(西北工业大学)
专题命中 VLA模型 :action model(title);vision-language-action(abstract);VLA(abstract);分类 cs.RO、cs.AI
Comments The first three authors contributed equally