arXiv cs.LGOctober 2, 2026
JEPA-Bisim: Learning Robust Visual Representations for Planning with Joint-Embedding Predictive World Models
Excerpt
arXiv:2602.18639v2 Announce Type: replace Abstract: World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recent latent predictive architectures (JEPAs), including the DINO world model (DINO-WM), display a degradation in test time robustness due to their sensitivity to ``slow features". These include visual variations such as background changes and distractors that are irrelev