机构
*
Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong(香港中文大学机械与自动化工程系)
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学计算机科学与工程系)
;
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Department of Mechanical and Automation Engineering, T Stone Robotics Institute, Shun Hing Institute of Advanced Engineering, Multi-Scale Medical Robotics Center, and Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学机械与自动化工程系、T Stone机器人研究所、Shun Hing先进工程研究所、多尺度医疗机器人中心以及医学智能与XR研究所)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
PixelVLA:推进视觉-语言-动作模型中的像素级理解
Wenqi Liang, Gan Sun, Yao He, Jiahua Dong, Suyan Dai, Ivan Laptev, Salman Khan, Yang Cong
机构
*
University of Trento(特伦托大学)
;
School of Automation Science and Engineering, South China University of Technology(华南理工大学自动化科学与工程学院)
;
Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学)
;
Australian National University(澳大利亚国立大学)
What Matters in Building Vision-Language-Action Models for Generalist Robots
在通用机器人中构建视觉-语言-动作模型所关注的关键因素
Xinghang Li, Peiyan Li, Long Qian, Minghuan Liu, Dong Wang, Jirong Liu, Bingyi Kang, Xiao Ma, Xinlong Wang, Di Guo, Tao Kong, Hanbo Zhang, Huaping Liu
机构
*
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
ByteDance Research(字节跳动研究院)
;
CASIA MAIS-NLPR
;
Shanghai Jiao Tong University(上海交通大学)
;
National University of Singapore(新加坡国立大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
CommentsThe Sigma model has been open-sourced on Hugging Face. Weights, dataset, some scripts, and logs are all available. The link is: https://huggingface.co/Veltraxor/Sigma
机构
*
TGAI Lab, School of Engineering, Westlake University(西拉库大学工程学院TGAI实验室)
;
Pennsylvania State University(宾夕法尼亚州立大学)
;
Sony Research, Sony(索尼研究实验室,索尼)
;
Xidian University(西安电子科技大学)
Developing Vision-Language-Action Model from Egocentric Videos
Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori
机构
*
Kyoto University(京都大学)
;
National Institute of Informatics(国家信息研究所)
;
Institute of Science Tokyo(东京科学研究所)
;
NII LLMC(日本信息机构语言模型中心)
;
Sony Interactive Entertainment(索尼互动娱乐)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Kun Wu, Zhiyuan Xu, Ning Liu, Ran Cheng, Chaomin Shen, Yaxin Peng, Feifei Feng, Jian Tang
机构
*
East China Normal University(东华师范大学)
;
Midea Group, AI Lab(美的集团人工智能实验室)
;
Syracuse University(雪城大学)
;
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
;
Shanghai University(上海大学)
Comments15 pages, 8 figures, 4 tables. In press in Acta Astronautica (Special Issue: 76th IAC). Presented at the 76th International Astronautical Congress (IAC 2025), Paper IAC-25-A4,1,6,x95796
The Zero-Age Massive Stellar Population of W49A from VLA Observations
基于VLA观测的W49A零龄大质量恒星群
M. Juárez-Gama, R. Galván-Madrid, G. Suárez, C. G. De Pree, J. J. Tobin, G. Bruzual, A. Ginsburg, J. E. Mendoza-Torres, H. B. Liu, D. Wilner, T. Nony, A. Parra-López, R. Rivera-Soto, A. F. McLeod, N. Cunningham, X. Lu, E. F. Jiménez-Andrade