Motivation
Standard JEPAs mask spatiotemporal patches and predict their latents, but a patch is not a meaningful unit of the world — distinct objects bleed into the same latent dimensions. That entanglement defeats the queries a world model most needs to answer: if I intervene on this one entity, what changes? Counterfactual reasoning and credit assignment both require knowing which part of the state corresponds to which thing. Causal-JEPA restructures the representation around objects, so that masking and prediction operate on entities rather than pixels, and the learned dynamics become modular and intervention-friendly.
How it works
The context encoder produces an object- or entity-centric set of latents — one slot per entity — rather than a flat grid of patch tokens. The self-supervised signal comes from masking at the level of objects: a whole entity's latent is hidden, and the predictor must infer it from the remaining entities and from the action. Training keeps the JEPA core — predict target-encoder latents under stop-gradient with anti-collapse — but the factored structure means the transition $\hat z_{t+1}=g_\phi(z_t,a_t)$ operates on disentangled slots.
Because each slot tracks one entity, the model isolates how that entity (and any intervention on it) propagates through the scene. This yields causal, modular dynamics and makes counterfactual edits — change one slot, predict the consequence — natural to express.
The objective
With slot set $\{z_t^{(i)}\}$ and a masked entity $m$, the loss predicts the masked slot's target latent from context slots and action:
$$\mathcal{L} = \big\lVert g_\phi\big(\{z_t^{(i)}\}_{i\neq m},\, a_t\big) - \operatorname{sg}\big[\bar f_\theta\big(o^{(m)}\big)\big] \big\rVert^2 + \lambda\,\mathcal{R}(Z).$$
The transition acts per slot, $\hat z_{t+1}^{(i)}=g_\phi(z_t,a_t)^{(i)}$, so reasoning over a sparse set of entities replaces dense patch prediction. Anti-collapse keeps the slots informative, and the object structure makes the cost of evaluating an intervention scale with the number of entities, not pixels.
Key results & what's novel
The novelty is object-level latent masking to obtain object-centric world models. Factoring the latent into causal entities supports counterfactual reasoning — silence or alter one slot and predict the rest — and makes planning more efficient, since search and intervention evaluation over a handful of entities is combinatorially cheaper and more interpretable than dense pixel-space rollouts. By aligning masking with the world's causal units, Causal-JEPA moves JEPA world models toward systematic generalisation and explicit intervention semantics rather than mere appearance prediction.
Strengths & limitations
- + Disentangled entity slots enable clean counterfactual queries.
- + Object-level reasoning makes planning and intervention search cheaper and more interpretable.
- + Modular dynamics aid systematic generalisation across scene configurations.
- − Requires reliable object/slot discovery, which is itself hard in cluttered or deformable scenes.
- − The number of slots and their binding to entities are design choices that may not match the true causal structure.
- − Object-centric encoders can struggle with stuff (backgrounds, fluids) that lacks discrete entities.