发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PhaseLoRA是面向连续动作VLA策略的轻量级PEFT方法,通过控制条件式LoRA实现阶段依赖适配,在LIBERO数据集上较匹配参数的高秩LoRA基线提升平均成功率12.2个百分点,性能优于其他LoRA变体。
AI 中文摘要
参数高效微调(PEFT)是适配预训练视觉-语言-动作(VLA)策略的自然方式,但多数适配器设计在整个控制滚动过程中应用时间静态更新,忽略了连续动作操作的阶段依赖特性。此类策略会经历不同阶段,包括接近、接触过渡、抓取、运输和放置,每个阶段需要不同的适配行为。我们提出PhaseLoRA,一种轻量级LoRA参数化方法,在每个动作块预测步骤使用两个弱监督描述符(精细控制倾向和事件/边界强度)来调节适配。PhaseLoRA在动作专家中调制LoRA左因子,使有效低秩更新方向随时间变化,同时保持主干网络基本冻结。在LIBERO数据集上,PhaseLoRA比参数匹配的高秩LoRA基线的平均成功率提升12.2个百分点,且优于更强的LoRA变体。消融实验表明,随机时间调制和标量门控无法复现完整模型的性能,而更新方向分析显示与预测控制描述符相关的结构化时间变化。这些结果确立了轨迹内条件调节作为连续动作VLA策略的有效轻量级PEFT维度。
英文摘要
Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs apply temporally static updates throughout a control rollout, overlooking the phase-dependent nature of continuous-action manipulation. Such policies traverse distinct regimes, including approach, contact transition, grasping, transport, and placement, each requiring different adaptation behaviors. We propose \textbf{PhaseLoRA}, a lightweight LoRA parameterization that conditions adaptation at each action-chunk prediction step using two weakly supervised descriptors: fine-control tendency and event/boundary intensity. PhaseLoRA modulates the LoRA left factor in the action expert, allowing the effective low-rank update direction to vary over time while keeping the backbone largely frozen. On LIBERO, PhaseLoRA improves average success rate by 12.2 points over a matched-parameter high-rank LoRA baseline and outperforms stronger LoRA variants. Ablations show that random temporal modulation and scalar gating do not reproduce the performance of the full model, while update-direction analyses reveal structured temporal variation associated with the predicted control descriptors. These results establish within-trajectory conditioning as an effective lightweight PEFT axis for continuous-action VLA policies.