发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SUNTA方法,通过解耦训练策略保留惊讶信号,并利用内部不一致性作为自上而下的惊讶度量在想象展开中确定分块边界,在2D和3D视频预测任务中优于基线,能在250步内保持准确预测。
AI 中文摘要
层次化状态空间模型(HSSMs)通过将序列分割为时间块,为长时域预测提供了一种有前景的方法。然而,其性能取决于如何确定分块边界。虽然先前的HSSMs通常依赖于固定长度分块或基于相似性的边界检测,但这些方法往往与数据的内在时间结构不一致。我们认为分块应由预测误差驱动,这更直接地指示何时需要更长范围的上下文。尽管如此,将基于惊讶的分块整合到HSSMs中引入了关键挑战,包括端到端训练期间的层次化崩溃以及在开环预测期间缺乏惊讶信号。为了解决这些问题,我们提出了基于惊讶的嵌套时间抽象(SUNTA),该方法采用解耦训练策略来保留惊讶信号,并使用内部不一致性作为自上而下的惊讶度量,在想象展开中确定分块边界。在2D和3D环境中的视频预测任务上的实验表明,SUNTA优于基线,独特地在250个时间步内保持准确预测,而所有基线在前10个时间步内就退化了。
英文摘要
Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks. However, their performance hinges on how chunk boundaries are determined. While prior HSSMs typically rely on fixed-length chunking or similarity-based boundary detection, these methods often misalign with the intrinsic temporal structure of the data. We argue that chunking should instead be driven by prediction errors, which more directly indicate when longer-range context becomes necessary. Nevertheless, integrating surprise-based chunking into HSSMs introduces critical challenges, including hierarchical collapse during end-to-end training and the absence of surprise signals during open-loop prediction. To address these issues, we propose Surprise-based Nested Temporal Abstraction (SUNTA), a method that employs a decoupled training strategy to preserve surprise signals and uses internal inconsistency as a top-down surprise metric to determine chunk boundaries within imagined rollouts. Experiments on video prediction tasks in 2D and 3D environments demonstrate that SUNTA outperforms baselines, uniquely maintaining accurate predictions over 250 timesteps, whereas all baselines degrade within the first 10 timesteps.