arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01049cs.AIcs.CVcs.LG

FactorJEPA:将整体未来分解为布局-智能体-交互通道以应对拥挤混乱的全球南方城市场景

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das

AI总结:

针对全球南方拥挤混乱的城市场景,该研究构建了首个大规模数据集DENSEWORLD-115k,并提出FactorJEPA模型,通过分解结构提升JEPA性能,相关结果在不同主干网络上具有一致性。

AI中文摘要:

世界模型因其捕捉和预测物理世界结构与动态的能力而备受关注,联合嵌入预测架构(JEPA)是其中极具前景的方向。本文研究了一个尚未充分探索的场景:人口稠密、拥挤且混乱的全球南方城市环境,我们将其命名为DENSEWORLD。与现有评估中占主导的低密度、车道结构化场景不同,这些场景呈现出模糊的空间边界、极端的智能体异质性、持续的遮挡以及混合交通下的快速社会协商。我们为该场景构建了首个大规模数据集:覆盖22个城市的1000小时驾车、步行及航拍视频。现有JEPA框架难以在异质性和部分可观测性下保留密集交互动态,我们提出FactorJEPA,将世界结构作为首要预测基元,而非在单一潜空间中编码未来,而是通过可见性门和分离子空间组合布局、实体与交互,以保留部分观测到的智能体并抑制跨因子捷径。FactorJEPA在四项指标上均有提升:(i)未来潜空间精度(Future-frame L1)、(ii)干预敏感预测(Causal L1)、(iii)视觉证据减少时的鲁棒性(Mask-ratio slope),同时揭示了(iv)可复现的运动-信息权衡(Motion cosine)。在2B和1B V-JEPA 2.1主干网络上,方法排名的相关系数ρ为0.895至0.978,结果一致。我们公开发布了DENSEWORLD-115k数据集(https URL)和经微调训练的FactorJEPA检查点(https URL)。

英文摘要:

World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction. We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we call DENSEWORLD. Unlike the lower-density, lane-structured settings that dominate existing evaluations, these scenes exhibit soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. We introduce the first large-scale dataset for this regime: 1,000 hours of drive-through, walk-through, and aerial video across 22 cities. Existing JEPA formulations struggle to preserve dense interaction dynamics under heterogeneity and partial observability. We introduce FactorJEPA, which makes world structure a first-class predictive primitive. Rather than encoding the future in a monolithic latent, it composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts. FactorJEPA improves (i) future-latent accuracy (Future-frame L1), (ii) intervention-sensitive prediction (Causal L1), and (iii) robustness to reduced visual evidence (Mask-ratio slope), while exposing (iv) a reproducible motion-information trade-off (Motion cosine). Method rankings replicate across 2B and 1B V-JEPA 2.1 backbones, with rho = 0.895 to 0.978. We publicly release the DENSEWORLD-115k dataset (https://huggingface.co/datasets/anonymousML123/denseworld-115k) and the surgery-trained FactorJEPA checkpoints (https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).

补充信息

↑