ST4VLA: Spatially Guided Training for Vision-Language-Action Models
ST4VLA:基于空间引导的视觉-语言-动作模型训练
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; The Hong Kong University of Science and Technology(香港科技大学) ; Southern University of Science and Technology(南方科技大学) ; Fudan University(复旦大学)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);grounding(abstract)
AI总结 ST4VLA通过空间引导训练提升视觉-语言-动作模型的性能,实现对机器人任务的更稳健和可泛化的学习。
Comments Spatially Training for VLA, Accepted by ICLR 2026