arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

中位数时间集成:面向动作分块视觉运动策略的无训练鲁棒聚合

Median Temporal Ensembling: Training-Free Robust Aggregation for Action-Chunked Visuomotor Policies

Yuhang Jiang

arXiv 2609.27167首次发表:更新:

发表机构

University of Trento(特伦托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对动作分块视觉运动策略,提出中位数时间集成,以坐标中位数替代指数加权平均,实现无训练鲁棒聚合,在对抗和故障场景下显著优于均值。

AI 中文摘要

动作分块的视觉运动策略预测重叠轨迹,因此每个执行的动作都涵盖在多个预测中。时间集成通过指数加权平均组合这些预测来平滑执行过程。一个被破坏的预测可以无界地移动聚合结果:其崩溃点为零。我们使用对抗性破坏来考验这一部署的聚合器,并比较两种保证类型。度量保证限制了给定大小扰动下的响应。组合保证则限制了在覆盖一个时间步的M个候选中至多q个被破坏时的损害,无论其大小如何。编码器对抗微调在已发布的补丁攻击下恢复了44%的损失,但当攻击者的步长增加后仅恢复7.3%。相比之下,同一候选集的坐标中位数在攻击优化增强时保持其恢复比例平稳。中位数时间集成只需一行代码且无需重新训练。在25种(配置,破坏级别)组合中,它从未比均值差,且在15种组合中显著更好。它还迁移到第二类策略,并在无攻击者参与的情况下(例如摄像头帧空白)的故障中恢复性能。其对干净数据的影响取决于配置,从-0.04到+0.07。我们还给出了边界:如果破坏使每个覆盖预测移动相同量,则整个统计族对此不可见,且没有等变聚合器能消除这种破坏。

英文摘要

Action-chunked visuomotor policies predict overlapping trajectories, so every executed action is covered by several predictions. Temporal ensembling smooths execution by combining these predictions with an exponentially weighted mean. One corrupted prediction can move the aggregate without bound: its breakdown point is 0. We use adversarial corruption to stress this deployed aggregator and to compare two kinds of guarantee. A metric guarantee bounds the response to a perturbation of a given size. A combinatorial guarantee instead bounds the damage when at most q of the M candidates covering a timestep are corrupted, whatever their size. Encoder adversarial fine-tuning recovers 44% of the loss under the published patch attack, but only 7.3% after the attacker's step size is increased. By contrast, the coordinate-wise median of the same candidate set keeps its recovered fraction flat as attack optimisation increases. Median temporal ensembling costs one line and requires no retraining. Across 25 (configuration, corruption-level) combinations it is never worse than the mean and is significantly better in 15. It also transfers to a second policy class, and it recovers performance under a failure with no attacker in the loop at all: camera frames that arrive blank. Its effect on clean data is configuration-dependent, from -0.04 to +0.07. We also give the boundary: corruption that shifts every covering prediction by the same amount is invisible to this whole family of statistics, and no equivariant aggregator can remove it.

Comments9 pages, 2 figures, 6 tables. Project page: https://avalon-s.github.io/MedianTE/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑