arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果历史测试时扩展用于自回归世界-动作模型中的失败恢复

Causal-History Test-Time Scaling for Failure Recovery in Autoregressive World-Action Models

Lin Li, Long Chen, Kwunhang, Wong, Jiaming Lei, Song Jin, Shucheng Du, Chuhan Zhang, Songchen Ma, Weihao Zhang, Jun Xiao, Kwang-Ting, Cheng

arXiv 2609.18016首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (HKUST); ACCESS – AI Chip Center for Emerging Smart Systems; Zhejiang University(香港科技大学; ACCESS——新兴智能系统AI芯片中心; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自回归世界-动作模型在操作失败时难以恢复的问题,提出无需训练的测试时扩展框架,通过进度感知触发、历史前缀恢复和假设验证三阶段,提升模拟与真实场景的任务成功率。

AI 中文摘要

世界-动作模型(WAMs)已成为机器人操作的一种有前景的范式,通过联合建模未来的视觉动态和机器人动作。然而,现有的WAMs主要在成功轨迹上训练,使得当真实世界执行偏离学习到的动态时,它们容易失败。这个问题在自回归WAMs中被放大,因为执行错误成为因果历史的一部分,并继续影响后续预测。为此,我们引入了\method{},一个无需训练的框架,将失败恢复重新表述为“因果历史上的测试时扩展”。该表述将恢复分解为三个耦合的决策:何时修订因果历史,何处恢复可靠的历史前缀,以及哪种历史配置最能支持后续执行。具体来说,\method{}通过三个阶段实现这些决策:1)**进度感知恢复触发器**检测持续的无进展,并仅在当前执行状态允许干预时触发恢复;2)**历史前缀恢复**识别不可靠的历史后缀,检索与当前物理状态匹配的历史锚点,并从保留的前缀重建因果KV状态,同时以最新的真实观察为条件;3)**假设验证**比较由完整历史、恢复前缀和完全重置假设引起的未来延续,并承诺最佳支持的假设。在模拟和真实世界操作设置中的实验表明,任务成功率持续提高,而消融研究确认了每个恢复阶段的贡献。

英文摘要

World-action models (WAMs) have emerged as a promising paradigm for robot manipulation by jointly modeling future visual dynamics and robot actions. However, existing WAMs are trained predominantly on successful trajectories, making them prone to failure when real-world execution diverges from the learned dynamics. This issue is amplified in autoregressive WAMs, where execution errors become part of the causal history and continue to influence subsequent predictions. To this end, we introduce \method{}, a training-free framework that reformulates failure recovery as \emph{test-time scaling over causal histories}. This formulation decomposes recovery into three coupled decisions: \emph{when} to revise the causal history, \emph{where} to recover a reliable history prefix, and \emph{which} history configuration best supports subsequent execution. Specifically, \method{} realizes these decisions through three stages: 1) \textbf{Progress-Aware Recovery Trigger} detects persistent non-progress and triggers recovery only when the current execution state permits intervention; 2) \textbf{History-Prefix Recovery} identifies the unreliable history suffix, retrieves a historical anchor matching the current physical state, and reconstructs the causal KV state from the retained prefix while conditioning on the latest real observation; and 3) \textbf{Hypothesis Verification} compares the future continuations induced by complete-history, recovered-prefix, and full-reset hypotheses, and commits the best-supported hypothesis. Experiments in both simulated and real-world manipulation settings demonstrate consistent improvements in task success, while ablations confirm the contribution of each recovery stage.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑