发表机构
The University of Manchester; Universität Hamburg; Comenius University Bratislava(曼彻斯特大学; 汉堡大学; 布拉迪斯拉发考门斯基大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出LAVLA框架用于VLA模型的潜在聚类分析,通过引入交叉注意力嵌入加权方法对GR00T N1.5模型分层研究,发现潜在聚类解耦时空与运动特征,提升了语言驱动机器人系统的可解释性。
AI 中文摘要
视觉-语言-动作(VLA)模型因能将语言与感知转化为动作,在机器人领域应用日益广泛,但驱动其行为的内部表征仍鲜为人知。我们提出LAVLA框架,用于VLA模型的潜在聚类分析,并对当前最先进的GR00T N1.5模型开展分层研究,重点关注其动作解码器。为更好地表征动作扩散过程中的潜在空间,我们引入一种基于交叉注意力的嵌入加权方法,该方法可放大相关特征,同时抑制信息度较低的特征。定量评估表明,加权聚类始终优于基线方法。为提升可解释性,我们为每个聚类提取了人类可理解的概念,将潜在表征与语义描述关联起来。我们的分析显示,潜在聚类逐步解耦时空与运动特征,中间层的表征愈发精细,到输出层趋于稳定。因此,LAVLA提升了语言驱动机器人系统的可解释性。
英文摘要
Vision-Language-Action (VLA) Models are increasingly used in robotics for their ability to ground language and perception into action, yet the internal representations driving their behaviour remain poorly understood. We propose LAVLA, a framework for latent cluster analysis of VLA models, and conduct a layer-wise study of the state-of-the-art GR00T N1.5 model, with particular focus on its action decoder. To better characterise the latent space during action diffusion, we introduce a cross-attention-based embedding-weighting method that amplifies relevant features while suppressing less informative ones. Quantitative evaluation shows that weighted clustering consistently outperforms the baseline. To improve interpretability, we extract human-interpretable concepts for each cluster, linking latent representations to semantic descriptions. Our analysis shows that latent clusters progressively disentangle spatiotemporal and kinematic features, with representations becoming more refined in the middle layers and stabilising toward the output. As such, LAVLA advances the interpretability of language-driven robotic systems.