arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

紧凑型机器人策略需要细粒度视觉表示

Compact Robot Policies Need Fine-Grained Visual Representations

Nanhe Chen, Runqiu Yang, Jiawei Tang, Sichao Liu, Yuquan Wang

arXiv 2610.08183首次发表:更新:

发表机构

X Square Robot; KTH(X Square Robot; KTH)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出紧凑策略CoRP,证明视觉表示是性能关键,预训练、任务适应和压缩缺一不可,以48.9M参数匹配大40-163倍系统的性能。

AI 中文摘要

多任务操作策略在架构、规模和预训练先验上同时存在差异,因此已发表的比较无法将性能归因于任何单一组件。我们认为性能差异主要来源于视觉表示,而参数规模和生成式先验在很大程度上是次要的。为验证这一观点,我们构建了CoRP(压缩表示策略),一种刻意保持紧凑的策略(48.9M参数,无视觉-语言模型,无视频生成先验),其分解为表示提取器和流匹配动作生成器。它在LIBERO上达到97.0%,在RoboTwin 2.0 Clean/Randomized上达到75.78%/73.36%,与规模大40.9-163.6倍的系统相当。固定动作生成器,我们逐一改变提取器属性。预训练初始化至关重要:随机初始化的ViT-S/14在LIBERO上降至78.1%,ImageNet ResNet-34降至74.5%。仅预训练不够,冻结编码器损失19.8个百分点。压缩同样重要:将每个视图重采样至48个token优于传递所有patch token(97.0% vs 83.2%),而对这些token使用变分信息瓶颈比硬token预算更差,将LIBERO-Goal从95.8%降至33.0%,因为抑制了策略依赖的指令相关token选择。语言条件仅在观察使目标模糊时起作用(LIBERO-Goal:9.2%提升至95.8%),而在RoboTwin 2.0中,观察明确,移除语言条件反而略微提升成功率。因此,我们主张紧凑策略在表示经过预训练、任务适应和压缩时有效。项目页面:此https URL。

英文摘要

Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental. To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator. It reaches 97.0% on LIBERO and 75.78%/73.36% on RoboTwin 2.0 Clean/Randomized, matching systems 40.9-163.6x larger. Holding the action generator fixed, we then vary one extractor property at a time. Pretrained initialization is decisive: a random ViT-S/14 drops to 78.1% and an ImageNet ResNet-34 to 74.5% on LIBERO. Pretraining alone is not enough, as freezing the encoder costs 19.8 points. Compression matters as much: resampling each view to 48 tokens beats passing all patch tokens (97.0% vs 83.2%), and a variational information bottleneck over those tokens is worse than a hard token budget, cutting LIBERO-Goal from 95.8% to 33.0% by suppressing the instruction-dependent token selection the policy relies on. Language conditioning contributes only where the observation leaves the goal ambiguous (LIBERO-Goal: 9.2% to 95.8%), while on RoboTwin 2.0, where observations are unambiguous, removing it slightly improves success. Therefore, we argue that a compact policy works when its representation is pretrained, task-adapted, and compressed. Project page: https://corp-policy.github.io/

Comments35 pages, 21 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑