发表机构
Ghent University(根特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究法证式复现ACT中CVAE编码器消融实验,发现原始成功率大幅下降未重现,且潜变量在推理时未使用、重建益处甚微,跳过编码器可提升训练吞吐量。
AI 中文摘要
动作分块Transformer(ACT)被广泛用于从示范中学习机器人操作。其条件变分自编码器包含一个编码器,旨在训练期间捕获示范之间的差异。原始ACT论文报告称,在具有人类示范的两个模拟任务上,移除编码器使平均成功率从35%降至2%。我们在原始代码中重新运行了这一消融实验,并检查了发现是否依赖于实现或训练数据。已发表的下降在我们的测试中未重现,尽管成功率的小幅提升或下降仍不确定。为调查差异,我们改变了训练时长和评估时检查点的选择方式。两者都能逆转哪个策略得分更高,但已发表下降的原因仍未知。仅凭成功率无法确定编码器是否提供了有助于策略重建示范动作的信息。在测试的ACT基准上,在潜信息惩罚的每个测试非零权重下,采样的潜变量几乎不提供重建益处。在推理时,ACT不使用此潜变量并将其设为零。在我们计时的两种实现中,跳过编码器都增加了训练吞吐量。我们发布了代码、评估工具和结果,以便他人重复比较并在其他任务上测试编码器。
英文摘要
Action Chunking Transformers (ACT) are widely used to learn robot manipulation from demonstrations. Their conditional variational autoencoder includes an encoder meant to capture differences between demonstrations during training. The original ACT paper reported that encoder removal dropped the mean success rate from 35% to 2% on two simulated tasks with human demonstrations. We re-ran this ablation in the original code and checked whether the findings depend on the implementation or training data. The published drop does not reappear in our tests, although smaller gains or losses in success rate remain uncertain. To investigate the discrepancy, we varied training length and how checkpoints are selected for evaluation. Both can reverse which policy scores higher, but the published drop's cause remains unknown. Success rates alone leave open whether the encoder provides information that helps the policy reconstruct demonstrated actions. On the tested ACT benchmark, the sampled latent provides little reconstruction benefit at every tested nonzero weight of the penalty on latent information. At inference, ACT leaves this latent unused and sets it to zero. Skipping the encoder increases training throughput in both implementations we timed. We release code, evaluation tools and results so others can repeat the comparisons and test the encoder on other tasks.