机构
*
Digital Environment Research Institute (DERI), Queen Mary University of London(伦敦大学玛丽女王学院数字环境研究所(DERI))
;
Shanghai Institute of Artificial Intelligence for Education, East China Normal University(华东师范大学上海人工智能教育研究院)
;
National Institute of Education, Nanyang Technological University(南洋理工大学国家教育学院)
;
School of Computer Science and Engineering, Jiangsu University of Science and Technology(江苏科技大学计算机科学与工程学院)
Comments8 pages, 5 figures. Introduces ForeTime-VLA, a causal future-token distillation method for conveyor-belt manipulation from a frozen world action model teacher
SkyNative: A Native Multimodal Architecture for Remote Sensing Vision-Language Understanding
SkyNative: 一种面向遥感视觉证据推理的原生多模态框架
Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan, Lang Sun, Jiaqi Liu, Jiashun Zhu, Jiasen Hu, Xu Na, Bo Yang
机构
*
College of Computer Science and Technology, Jilin University, China(吉林大学计算机科学与技术学院)
;
Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education(教育部符号计算与知识工程重点实验室)
Comments19 pages (10 pages main text + appendix), 10 figures. Accepted at the 34th ACM International Conference on Multimedia (ACM MM 2026), Rio de Janeiro, Brazil, November 10--14, 2026
机构
*
School of Advanced Interdisciplinary Sciences, UCAS(UCAS交叉学科研究院)
;
School of Electronic, Electrical and Communication Engineering, UCAS(UCAS电子电气与通信工程学院)
;
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(中国科学院计算技术研究所人工智能安全国家重点实验室)
;
Alibaba Group(阿里巴巴集团)
;
School of Computer Science and Technology, UCAS(UCAS计算机科学与技术学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Key Laboratory of Big Data Mining and Knowledge Management, UCAS(UCAS大数据挖掘与知识管理重点实验室)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.LG