AI 中文总结
SJEPA是一种无重构的JEPA框架,结合符号定律与神经修正学习简洁潜在动力学,在受控摆实验中展现更优性能,揭示了多目标间的可控权衡。
AI 中文摘要
联合嵌入预测架构(JEPA)通过从上下文嵌入预测目标嵌入来学习抽象状态,但其转移模型通常是不透明的神经映射。我们引入SJEPA,这是一种无重构的JEPA框架,可学习其诱导动力学允许紧凑符号描述的预测表示。其混合转移结合了符号定律和正则化神经修正,用于所选语法之外的动力学。核心原则是学习最简单的充足动力学:表示约束保留信息丰富、未崩溃的预测坐标,而算子压缩则有利于保持预测充足性的低复杂度符号-神经转移。我们通过诱导动力学复杂度形式化该原则,分析预测坐标的不可识别性,并表明无约束的算子压缩会产生通往表示崩溃的直接捷径。该框架支持交替的表示-方程学习和适配固定表示的符号动力学。在受控摆实验中,联合学习相比事后拟合发现了显著更简单的符号动力学,且具有更低的长程滚动误差和发散性,而无约束的单步诊断则实现了预测的崩溃捷径。在语法误配下,修正正则化保留了可表示的符号机制,并引导神经组件处理残差动力学。结果揭示了预测保真度、表示质量、符号简洁性和符号-神经分配之间可控制的权衡。
英文摘要
Joint-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings, but their transition models are typically opaque neural maps. We introduce SJEPA, a reconstruction-free JEPA framework that learns predictive representations whose induced dynamics admit compact symbolic descriptions. Its hybrid transition combines a symbolic law with a regularised neural correction for dynamics outside the selected grammar. The central principle is to learn the simplest adequate dynamics: representation constraints preserve informative, non-collapsed predictive coordinates, while operator compression favours low-complexity symbolic-neural transitions that remain predictively adequate. We formalise this principle through induced-dynamics complexity, analyse predictive-coordinate non-identifiability, and show that unconstrained operator compression creates a direct shortcut to representation collapse. The framework supports both alternating representation-equation learning and symbolic dynamics fitted to fixed representations. In controlled pendulum experiments, joint learning discovers substantially simpler symbolic dynamics with lower long-horizon rollout error and divergence than post-hoc fitting, while an unconstrained one-step diagnostic realises the predicted collapse shortcut. Under grammar misspecification, correction regularisation preserves the representable symbolic mechanism and directs the neural component towards residual dynamics. The results expose a controllable trade-off among predictive fidelity, representation quality, symbolic parsimony, and symbolic-neural allocation.
Comments42 pages