发表机构
Boston University; Shanghai Jiao Tong University(波士顿大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SR-JEPA,一种面向场景级点云的原生点JEPA,通过仅使用自包含的3D EMA目标训练,在3D场景实体缺失时可查询预测隐状态,在ARKitScenes和Sr3D数据集上取得了良好的语义预测性能,揭示了可组合的3D预测状态。
AI 中文摘要
联合嵌入预测架构(JEPAs)通过预测缺失观测的隐表示进行学习,但许多掩码JEPAs的评估主要基于它们生成的编码器。本文探究了当原生3D场景中完全缺失某个实体时,训练后的预测通路本身会推断出什么。我们提出SR-JEPA,一种面向场景级点云的原生点JEPA,其原始冻结的预测通路可在指定位置查询。评估时,先移除一个物体的所有点再进行编码,并在其质心处替换为相同的无形状32点查询。训练仅使用自包含的3D EMA目标,不使用重建、语义标签、语言或提升的2D特征。在5953个保留的ARKitScenes物体上,估算的隐状态达到43.13%的语义类别宏准确率,比最强基线高22.18个百分点;随机化预测通路会使准确率下降9.78个百分点,替换为匹配的 donor 上下文则下降21.98个百分点。在8570个Sr3D支持对上,完整隐状态达到41.15 AP;从缺失物体隐状态解码出的类别,结合锚点类别与几何信息,达到39.37 AP,剩余1.78个百分点的未解决残差。这些结果揭示了一种可查询、可组合的3D预测状态:模型能完成依赖上下文的实体内容,下游计算可将其与度量几何结合。
英文摘要
Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce. We ask what a trained predictive pathway itself infers when an entire entity is absent from a native 3D scene. We introduce SR-JEPA, a point-native JEPA for scene-scale point clouds whose original frozen predictive pathway can be queried at a supplied location. At evaluation, every point of one object is removed before encoding and replaced by the same shape-free 32-point query at its centroid. Training uses only self-contained 3D EMA targets: no reconstruction, semantic labels, language, or lifted 2D features. On 5,953 held-out ARKitScenes objects, the imputed latent reaches 43.13% semantic-identity macro accuracy, 22.18 points above the strongest floor. Randomizing the prediction path removes 9.78 points, while substituting matched donor context removes 21.98 points. On 8,570 Sr3D support pairs, the full latent reaches 41.15 AP; identity decoded from the missing-object latent, combined with anchor identity and geometry, reaches 39.37 AP, leaving an unresolved 1.78-point residual. These results reveal a queryable, compositional 3D predictive state: the model completes context-dependent entity content, which downstream computation combines with metric geometry.
Comments17 pages, 5 figures, 9 tables