arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测量动作分块背后的稳定性假设

Measuring the Stability Assumption Behind Action Chunking

Aryan Goyal

arXiv 2610.01626首次发表:更新:

发表机构

Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过注入误差测量动作分块中误差传播速率,发现稳定状态罕见且误差放大常见,表明误差累积论点不完整,需显式训练闭环反应性。

AI 中文摘要

动作分块通过行为克隆提升了策略的性能,已有多种机制被提出来解释其原因,包括时间一致性、视野缩减、表示学习以及误差累积的减少。我们转而研究动作误差一旦进入系统后会发生什么。在每个状态下,我们注入一个小的动作误差,并测量在两种执行机制下其增长或缩小的速度:开环,即分块的其余部分在不重新规划的情况下重放;闭环,即策略在扰动后重新规划。拟合的速率将每个状态标记为收缩、扩张或未确定。在来自三个基准套件的十二个操作任务中,我们发现自信稳定的状态很少见,而在传播速率可解析的状态中,误差放大很常见。我们进一步发现,测得的传播速率强烈依赖于拟合视野:放大通常集中在前期,因此短窗口可能高估更长视野的传播。最后,我们基于这些标签训练预测器,发现仅从相机帧和本体感觉就能恢复状态的开环机制,而其闭环传播只能部分恢复,因为它还取决于策略在扰动后的行为。这些结果表明,仅凭误差累积的论点并不能完整解释动作分块:无论是被动的开环动力学还是策略重新规划都不能一致地收缩注入的误差,而且重新规划很少将开环放大转变为自信的收缩。这表明闭环反应性应被显式训练,使用面向扰动和树覆盖的训练来让策略暴露于必须恢复的偏差,而不是期望从标准模仿学习中可靠地涌现。

英文摘要

Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes: open-loop, where the rest of the chunk is replayed without replanning, and closed-loop, where the policy replans after the perturbation. The fitted rate labels each state as contracting, expanding, or unresolved. Across twelve manipulation tasks from three benchmark suites, we find that confidently stable states are rare, while error amplification is common among states whose propagation rate can be resolved. We further find that the measured propagation rate depends strongly on the fitting horizon: amplification is typically front-loaded, so short windows can overestimate longer-horizon propagation. Finally, we train predictors on these labels and find that a state's open-loop regime can be recovered from camera frames and proprioception alone, while its closed-loop propagation is only partially recoverable because it also depends on how the policy acts after the perturbation. These results suggest that error-compounding arguments alone do not provide a complete account of action chunking: neither passive open-loop dynamics nor policy replanning consistently contracts an injected error, and replanning rarely turns open-loop amplification into confident contraction. This suggests that closed-loop reactivity should be trained explicitly, using perturbation- and tree-coverage-oriented training to expose policies to deviations they must recover from, rather than expected to emerge reliably from standard imitation learning.

Comments18 pages, 9 figures, 18 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑