arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30996cs.LGcs.AI

视觉-语言-动作模型的线性表示假说

The Linear Representation Hypothesis for Vision-Language-Action Models

Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出视觉-语言-动作模型的线性表示假说理论,统一表示与策略,证明线性探针可恢复物理量演化,并引入签名广义线性模型实现线性引导,在导航实验中验证。

中文摘要 AI 辅助

线性表示假说(LRH)已成为通过大型语言模型(LLMs)的内部表示来测量和干预语义信息的标准视角。越来越多的研究开始将这一视角扩展到视觉-语言-动作(VLA)模型,但具身交互的动态特性带来了额外的挑战。与LLMs中常研究的语义属性(如性别或语言)不同,VLA中的物理量(QoI)随系统动力学共同演化:表示影响策略选择的动作,动作改变物理状态,进而影响下一个表示。在本文中,我们为VLA提出了一个基于签名(signature)的理论化LRH公式,统一了表示和策略。在表示方面,我们证明了存在这样的表示,使得在候选动作轨迹下,QoI的未来演化可以通过线性探针(linear probing)恢复。在策略方面,我们引入了用于随机动作块的签名广义线性模型。该结构在自然参数空间中沿线性路径产生期望未来QoI的单调变化,从而实现线性引导(linear steering)。我们在一个平面控制仿射导航实验中构建了显式的oracle表示,并验证了预测的线性探针和引导机制。

英文摘要

The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in a VLA evolves jointly with the system dynamics: the representation influences the actions selected by the policy, which alter the physical state and, in turn, the next representation. In this paper, we develop a theoretical, signature-based formulation of the LRH for VLA that unifies representations and policies. On the representation side, we establish the existence of representations from which the future evolution of a QoI under a candidate action trajectory can be recovered via linear probing. On the policy side, we introduce a signature generalized linear model for stochastic action chunks. This structure yields a monotonic change in the expected future QoI along linear paths in natural parameter space, enabling linear steering. We first validate this structure in a controlled oracle setting, then examine whether the same probing and steering mechanisms emerge in a pretrained VLA.

发表机构

  • KAIST(韩国科学技术院)
  • University of Illinois Urbana--Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑