arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11959cs.LGcs.CL

空间作为干预不变量:分层城市与仿真空间智能的跨模态预测几何

Space as an Interventional Invariant: Cross-Modal Predictive Geometry for Stratified Cities and Em-Spaced Intelligence

  • Tsinghua University(清华大学)
  • University College London(伦敦大学学院)
  • Cardiff University(卡迪夫大学)

机构由 AI 辅助整理,请以论文原文为准。

Tao Yang, Xuhui Lin, Kunyao Li, Haijiang Li

AI总结:

本文提出将空间定义为干预不变量,构建跨模态预测几何框架,在理论上证明潜在空间可识别性,并扩展到分层城市系统,为空间认知与具身智能提供统一基础。

AI中文摘要:

空间是数学、物理学、空间认知、城市科学与具身智能中的基础概念,然而这些领域往往将空间结构视为共享的几何容器或一组不相关的表征。此类方法难以解释异质的感官与城市过程如何共同揭示一种共同的空间结构,尤其是在不同模态不共享相同度量或表征的情况下。本文通过将空间定义为干预不变量来填补这一空白:即在可行动作下保持局部兼容性与未来观测条件法则的最小关系结构。我们发展了一种跨模态预测几何,整合了局部状态空间、模态特定观测映射、动作群胚与规范预测状态商,并给出了识别干预性而非仅观测性结构的显式因果条件。关键理论结果表明,在联合点分离、等变性与干预忠实性条件下,潜在空间可识别至干预群的中心化子,从而将表征歧义缩减为残余坐标自由度。该框架进一步通过层值表征扩展到分层城市系统,使几何、物理、移动性、社会与经济层能够共存而不被归结为单一度量。在噪声下的合成实验评估了等变性、预测充分性、和乐、限制映射恢复、跨尺度一致性与上下文饱和性。所提出的框架为空间认知、城市科学、具身AI与仿真空间智能提供了统一且可证伪的基础。

英文摘要:

Space is a foundational concept across mathematics, physics, spatial cognition, urban science, and embodied intelligence, yet these fields often treat spatial structure either as a shared geometric container or as a collection of disconnected representations. Such approaches struggle to explain how heterogeneous sensory and urban processes can jointly reveal a common spatial structure, particularly when different modalities do not share the same metric or representation. This paper addresses this gap by defining space as an interventional invariant: the minimal relational structure that preserves local compatibility and the conditional laws of future observations under admissible actions. We develop a cross-modal predictive geometry that integrates local state spaces, modality-specific observation maps, an action groupoid, and a canonical predictive-state quotient, with explicit causal conditions for identifying interventional rather than merely observational structure. The key theoretical result shows that, under joint point separation, equivariance, and interventional faithfulness, the latent space is identifiable up to the centraliser of the intervention group, thereby reducing representational ambiguity to residual coordinate freedom. The framework is further extended to stratified urban systems using sheaf-valued representations, allowing geometric, physical, mobility, social, and economic layers to coexist without being reduced to a single metric. Synthetic experiments under noise evaluate equivariance, predictive sufficiency, holonomy, restriction-map recovery, cross-scale consistency, and context saturation. The resulting framework provides a unified and falsifiable foundation for spatial cognition, urban science, embodied AI, and em-spaced intelligence.

↑