VLA-0: Building State-of-the-Art VLAs with Zero Modification
机构 * NVIDIA
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * NVIDIA
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.AI
机构 * Intern Robotics Shanghai AI Laboratory(Intern Robotics上海AI实验室)
专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.CV、cs.AI
Comments Technical report
机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) ; Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所) ; New Laboratory of Pattern Recognition (NLPR)(模式识别新实验室) ; State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS)(多模态人工智能系统国家重点实验室) ; Beihang University(北航) ; Chinese University of Hong Kong(香港大学)
专题命中 GUI与屏幕智能体 :grounding(abstract)
Comments Demo Page: https://embodiedcoder.github.io/EmbodiedCoder/