arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PredVLA:用于机器人操作的亚百万级参数预测编码策略

PredVLA: Predictive Sensorimotor Modeling for Sub-Million-Parameter Robot Manipulation

Hiroki Sawada, Shunichi Kasahara

arXiv 2608.26673首次发表:更新:

发表机构

Sony Computer Science Laboratory(索尼计算机科学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PredVLA是仅0.68M参数的无机器人数据预训练预测编码策略,在LIBERO基准上远超参数匹配的Transformer、LSTM策略,实现强语言条件机器人操作性能且具备显式在线修正机制。

AI 中文摘要

大型预训练视觉-语言-动作模型主导了现代机器人操作基准测试,但目前尚不清楚实现强语言条件控制所需的模型规模,或完全不同的控制架构在参数预算小得多的情况下能否保持竞争力。我们提出PredVLA,这是一种语言条件预测编码策略,仅包含0.68百万个可训练网络参数,且未使用机器人数据进行预训练;其分层生成式循环动力学可预测视觉特征和本体感觉,而观测结果仅通过对所得感官预测误差的在线推理影响隐状态。在LIBERO基准测试中,PredVLA在三个短 horizon 套件上的平均成功率达86.9%,若纳入长 horizon 套件则为75.4%。在使用相同冻结前端、演示数据、动作解码器和评估协议的受控对比中,PredVLA的平均成功率分别是参数匹配的Transformer策略的3.7倍、LSTM策略的7.4倍。预测编码公式还使观测驱动修正的贡献可直接测量:由于观测仅通过基于预测误差的隐推理影响循环状态,禁用该推理会得到精确的开环控制条件。这些结果共同表明,亚百万级参数的循环生成策略可在现代语言条件操作基准上实现强性能,同时为预测误差驱动的在线状态修正提供了显式机制。

英文摘要

Large pretrained vision-language-action models achieve strong robot-manipulation performance, while compact alternatives have largely pursued efficiency by compressing the prevailing observation-to-action paradigm. We investigate whether predictive sensorimotor modeling can make more effective use of a limited parameter budget than direct observation-to-action mapping. We present PredVLA, a language-conditioned predictive-coding policy with only 0.68 million trainable network parameters and no robot-data pretraining. Its hierarchical recurrent dynamics predict visual features and proprioception, while observations influence latent state only through prediction-error-driven online inference. On LIBERO, PredVLA achieves an 86.9% mean success rate across the three short-horizon suites and 75.4% across all four suites. Under a controlled comparison using the same frozen front end, demonstrations, action decoder, and evaluation protocol, PredVLA achieves 3.7x and 7.4x the three-suite mean success rates of parameter-matched Transformer and LSTM behavior-cloning policies, respectively. A mechanism-by-mechanism transition to the recurrent behavior-cloning baseline shows that replacing the predictive pathway with direct observation input produces the largest single performance drop, accounting for approximately $70\%$ of the endpoint gap. Further ablations identify distinct contributions from training-time latent inference, test-time error regression, hierarchical timescales, and sensory prediction-error channels. Together, these results support predictive sensorimotor modeling as a strong inductive bias for compact language-conditioned robot control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑