发表机构
Southwest University; Zhejiang University; Tianfu Securities Co., Ltd.(西南大学; 浙江大学; 天府证券股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对自回归视觉-语言-动作模型,通过对比分析式与数据驱动式动作分词化方法,发现重建误差无法可靠选择动作表示,提出需联合评估几何保真度、序列可预测性、解码器稳定性与闭环性能。
AI 中文摘要
离散动作分词化是自回归视觉-语言-动作(VLA)模型的核心,然而动作表示通常主要通过重建保真度来评估。我们通过在一个统一的分词化接口下比较固定的分析式、数据驱动的线性以及非线性神经表示,来探究哪些表示属性对闭环控制真正重要。在率失真分析、序列建模诊断以及3,500次LIBERO rollout中,表示排名随评估标准的变化而变化。PCA实现了比Temporal-DCT更低的标称重建误差,但在三个策略训练种子中,其产生的token序列可预测性较低,且平均已见任务成功率低3.0个百分点,其中在一个种子中策略排序发生反转。在匹配的种子42消融实验中,自编码器进一步降低了重建误差,但并未产生最强的策略,并且对离散token扰动表现出更高的敏感性。这些发现表明,仅凭重建保真度无法可靠地为自回归控制选择动作表示,从而促使对几何保真度、序列可预测性、解码器稳定性以及闭环性能进行联合评估。
英文摘要
Discrete action tokenization is central to autoregressive vision-language-action (VLA) models, yet action representations are often evaluated primarily through reconstruction fidelity. We ask which representation properties actually matter for closed-loop control by comparing fixed analytical, data-driven linear, and nonlinear neural representations under a unified tokenization interface. Across rate-distortion analysis, sequence-modeling diagnostics, and 3,500 LIBERO rollouts, representation rankings change with the evaluation criterion. PCA achieves lower nominal reconstruction error than Temporal-DCT, but produces less predictable token sequences and 3.0 percentage points lower mean seen-task success across three policy-training seeds, with the policy ordering reversing in one seed. In a matched seed-42 ablation, an autoencoder further reduces reconstruction error yet does not yield the strongest policy and exhibits greater sensitivity to discrete token perturbations. These findings show that reconstruction fidelity alone cannot reliably select action representations for autoregressive control, motivating joint evaluation of geometric fidelity, sequence predictability, decoder stability, and closed-loop performance.
Comments5 pages, 1 figure, 4 tables