arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23048cs.ROcs.LG

闭环崩溃的解剖:一个压缩VLA策略的因果案例研究

Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy

Fengze Jia

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过因果分析揭示压缩VLA策略在闭环执行中崩溃的结构化特征,并证明离线评估不足以作为验收测试,而少量闭环试验即可发现问题。

中文摘要 AI 辅助

压缩操作策略可能通过离线评估,但在闭环执行中失败;这种分离已在先前工作中确立,并非我们的主张。我们贡献了一个自然发生案例的因果解剖。对Octo-Base的8层蒸馏保留了86%的参数,通过了我们应用的所有离线检查(在家族自身的验证指标上教师比率为0.996和1.000),但在闭环中崩溃:在模拟WidowX拾取放置任务中,学生为0/72,而教师为40/72。崩溃是结构化的,而非弥散的:早期任务阶段逐渐退化(学生以教师90%的速率移动物体,并以55%的速率抓取),而运输到目标则完全失败,在所有训练变体中均为0%。配对的动作轨迹取证隔离了特征:一个负的、后期加重的$z$残差,大约是修复后量级的10倍,并且在基础蒸馏和两个持续分支中持续存在。四种标准疗法在匹配对照下失败:持续训练和域内离线数据使成功率保持为零,尽管后者可测量地改善了边际动作统计;命令级补偿在任何偏移下都未恢复任何效果,尽管相同的扰动会降低健康策略的性能;在命令通道中钳制症状保留了抓取能力,但成功率仍处于最低水平。一个最小对干预,用部署分布教师回放替换一半的训练流,并保持所有其他设置不变,恢复了与教师的同等性能(保留集上18/36对比17/36),消除了该特征,并恢复了类似教师的扰动响应曲线。我们主张存在性,而非普遍性。在操作上,离线门控,包括家族自身的验证指标,不足以作为压缩策略的验收测试;几十次闭环试验就足以发现它们遗漏的问题。

英文摘要

Compressed manipulation policies can pass offline evaluation while failing in closed-loop execution; this dissociation is established in prior work and is not our claim. We contribute a causal anatomy of one naturally occurring case. An 8-layer distillation of Octo-Base retains 86% of parameters, passes every offline check we applied (0.996 and 1.000 teacher-ratios on the family's own validation metrics), and collapses in closed loop: 0/72 vs. the teacher's 40/72 on a simulated WidowX pick-and-place task. The collapse is structured, not diffuse: early task stages degrade gradually (the student moves the object at 90% of the teacher's rate and grasps at 55%), while transport-to-target fails categorically, at 0% in every training variant. Paired action-trace forensics isolate the signature: a negative, late-heavy $z$ residual, roughly 10x its post-repair magnitude, and persistent across the base distillation and both continuation branches. Four standard therapies fail under matched controls: continued training and in-domain offline data leave success at zero, even though the latter measurably improves marginal action statistics; command-level compensation recovers nothing at any offset, although the same perturbations degrade healthy policies; clamping the symptom in the command channel preserves grasping, yet success stays at floor. A minimal-pair intervention that substitutes half of the training stream with deployment-distribution teacher rollouts, with every other setting held fixed, restores parity with the teacher (18/36 vs. 17/36 held-out), eliminates that signature, and recovers a teacher-like perturbation-response profile. We claim existence, not universality. Operationally, offline gates, including a family's own validation metrics, are insufficient acceptance tests for compressed policies; a few dozen closed-loop trials sufficed to find what they missed.

发表机构

  • The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑