arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34982cs.CVcs.RO

ActionUNet:利用高效多尺度微调提升VLA模型的鲁棒性

ActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuning

Di Zhu, Ziheng Yan, Fang Wan

首次发表
浏览论文内容

中文总结 AI 辅助

提出ActionUNet,一种高效多尺度微调框架,通过时间U-Net融合结构先验和条件SIREN连续解码器,提升VLA模型在杂乱环境中的鲁棒性,在多个基准上显著提高成功率。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型通过将多模态语义映射到物理动作,在机器人操作领域展现出巨大潜力。然而,这种映射本质上难以将粗粒度的语义与细粒度的时间执行对齐,导致VLA模型在杂乱环境中泛化能力有限且鲁棒性不足。为解决这一问题,我们提出了ActionUNet,一种高效的多尺度微调框架,以极小的计算成本增强预训练的VLA模型。ActionUNet首先在时间对齐的动作特征空间中构建一个轻量级的时间U-Net,以融合层次结构先验,有效弥合语义与时间执行之间的尺度差距。鉴于多尺度建模可能破坏微观时间连续性并引起机械振荡,ActionUNet随后采用条件SIREN作为连续动作解码器,该解码器配备显式的二阶平滑约束,保证时间连续性并减少高频运动抖动。通过平滑多尺度融合带来的时间不连续性,这种连续公式化方法减少了机械执行失败,同时保留了基础VLA模型的泛化能力和操作鲁棒性。在RoboTwin 2.0和LIBERO-Plus基准上的大量实验,以及真实世界的困难评估表明,ActionUNet分别将π0.5成功率绝对提升了9.8%、6.1%和11.4%,同时还能泛化到基于回归的OpenVLA-OFT骨干网络,凸显了其作为微调策略的有效性和高效性。代码和实现细节可在该https URL获取。

英文摘要

Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited generalization and insufficient robustness in cluttered environments. To overcome this issue, we propose ActionUNet, an efficient multi-scale fine-tuning framework that enhances pre-trained VLA models with minimal computational cost. ActionUNet first constructs a lightweight temporal U-Net within the temporal-aligned action feature space to fuse hierarchical structural priors, effectively bridging the scale gap between semantics and temporal executions. Recognizing that multi-scale modeling can disrupt microscopic temporal continuity and cause mechanical oscillations, ActionUNet then employs a conditional SIREN as a continuous action decoder. Equipped with explicit second-order smoothness constraints, this decoder guarantees temporal continuity and reduces high-frequency motion jitter. By smoothing temporal discontinuities from multi-scale fusion, this continuous formulation reduces mechanical execution failures while preserving the base VLA model's generalization and manipulation robustness. Extensive experiments on RoboTwin 2.0 and LIBERO-Plus benchmarks, together with real-world hard evaluations, demonstrate that ActionUNet significantly improves π0.5 success rates by absolute 9.8%, 6.1%, and 11.4%, respectively, while also generalizing to the regression-based OpenVLA-OFT backbone, highlighting its effectiveness and efficiency as a fine-tuning strategy. Code and implementation details are available at https://github.com/Di-Zhu123/ActionUNet.

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

↑