arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06823cs.CV

通用开放世界时间感知

Generalist Open-World Temporal Perception

  • Google DeepMind(谷歌DeepMind)
  • Lund University(隆德大学)
  • IMAR(罗马尼亚科学院数学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Cristian Sminchisescu

中文总结 AI 辅助

本文提出通用开放世界时间感知架构(GOWTPA),将感知与合成作为共享生成基底的不同条件化方式,以支持多模态智能与物理AI的基础层。

中文摘要 AI 辅助

下一代人工智能系统很可能在输入和输出上都是原生的时间性和多模态的:能够通过共享的世界表征进行对话、感知、预测、推理和合成。实现这一点需要一个时间感知基底,将感官流、语言和结构化输出整合在多模态世界模型中。我们寻求一种通用开放世界感知系统,将生物形态、自然物理结构和人工制品及其交互表示为连贯的、时间上持续的过程。该模型应从原始多模态流中推断几何、关节、语义、交互结构和不确定性;在遮挡和视角变化下保持身份;跨物种、形态、机制和材料进行泛化;并在遇到未知时弃权(不执行)或扩展其本体。目标是支持理解、预测、反事实推理和可控合成的结构化世界状态。近期研究表明,一些跨模态和类似推理的能力可以从大规模生成式视频预训练中涌现,类似于语言模型的扩展。然而,这些能力通常通过语言探针访问或通过逼真视频表达,使得显式的语义、几何或时间结构在很大程度上未被揭示。本文阐述了一种替代且互补的范式:在共享的通用开放世界时间感知架构(GOWTPA)中,感知和合成作为不同的条件化方式。识别、结构化预测和模拟作为同一生成基底的不同条件化方式出现,而推理和具身特定策略则基于所产生的世界状态。这将通用时间感知定位为更广泛的多模态智能和物理AI的潜在基础层。

英文摘要

The next generation of artificial intelligence systems will likely be natively temporal and multimodal in both inputs and outputs: able to converse, perceive, predict, reason, and synthesize through a shared world representation. Realizing this requires a temporal perceptual substrate integrating sensory streams, language, and structured outputs within a multimodal world model. We seek a generalist open-world perceptual system that represents biological forms, natural physical structures, and artifacts, and their interactions, as a coherent, temporally persistent process. The model should infer geometry, articulation, semantics, interaction structure, and uncertainty from raw multimodal streams; maintain identity through occlusion and viewpoint change; generalize across species, forms, mechanisms, and materials; and abstain or expand its ontology when encountering the unknown. The objective is a structured world state supporting understanding, prediction, counterfactual reasoning, and controllable synthesis. Recent work suggests that some cross-modal and reasoning-like capabilities can emerge from large-scale generative video pretraining, reminiscent of language-model scaling. Yet these capabilities are often accessed through language probes or expressed through photorealistic video, leaving explicit semantic, geometric, or temporal structure largely unexposed. This paper articulates an alternative and complementary paradigm: perception and synthesis as distinct conditionings within a shared Generalist Open-World Temporal Perception Architecture (GOWTPA). Recognition, structured prediction, and simulation arise as different conditionings of the same generative substrate, while reasoning and embodiment-specific policies build upon the resulting world state. This positions generalist temporal perception as a potential foundation layer for broader multimodal intelligence and physical AI.

补充信息

↑