AI 中文总结
研究如何从部分观测预测场景演变,提出GARFIELD概率模型,能学习潜在表示实现轨迹联合采样和运动分布访问,实验显示其在运动规划性能、采样速度及密度估计上优势显著。
AI 中文摘要
从部分观测预测场景如何演变需要考虑多种可能的未来,而非局限于单一轨迹。现有方法要么生成外观主导的视频预测,要么采样少量轨迹而未明确建模可能运动的分布。我们引入未来运动学潜在分布的目标感知表示(GARFIELD),这是一种场景运动学概率模型,能根据图像和时空稀疏约束学习未来可能分布的结构化时空潜在表示。同一潜在表示可实现所有轨迹的联合采样,并通过高效确定性密度解码器直接访问潜在运动分布。实验表明,该方法在运动规划性能上与大型视频生成模型相当,采样轨迹速度快97倍,估计运动密度比蒙特卡洛采样快两个数量级,可实现交互式探索和不确定性感知规划。
英文摘要
Predicting how a scene may evolve from partial observations requires reasoning about multiple possible futures rather than committing to a single trajectory. Existing approaches either generate appearance-dominated video predictions or sample a small number of trajectories without explicitly modeling the distribution of possible motion. We introduce Goal-Aware Representations of Future kInEmatic Latent Distributions (GARFIELD), a probabilistic model of scene kinematics that learns a structured spatio-temporal latent representation of the distribution over possible futures given an image and optional spatio-temporally sparse constraints. The same latent representation enables both joint sampling of all trajectories and direct access to the underlying motion distribution through an efficient deterministic density decoder. As a result, uncertainty about future motion can be localized to specific scene elements and timesteps and progressively refined through additional constraints. Experiments demonstrate strong motion planning performance competitive with large video generation models while sampling trajectories $97\times$ faster. Our method further estimates motion densities two orders of magnitude faster than Monte-Carlo sampling from motion generation models, enabling interactive exploration and uncertainty-aware planning.
CommentsAccepted at ECCV 2026. Project page: https://compvis.github.io/schroedingers_cat