At a glance
ProblemPatch-level masking entangles distinct objects in a single latent, making counterfactual reasoning ("what if only this entity changes?") and credit assignment hard.
Key ideaEncode the scene as a set of object/entity latents and mask at the level of entities, so the predictor learns modular, causal dynamics over slots.
ModalityVision (object-centric, + actions)
Target / maskingObject-level latent masking — predict a masked entity's latent from the others and from actions, against a target encoder.
Builds onJEPA latent prediction and object-centric / slot representations.
Used forCounterfactual reasoning, intervention semantics, and efficient object-level planning.

Motivation

Standard JEPAs mask spatiotemporal patches and predict their latents, but a patch is not a meaningful unit of the world — distinct objects bleed into the same latent dimensions. That entanglement defeats the queries a world model most needs to answer: if I intervene on this one entity, what changes? Counterfactual reasoning and credit assignment both require knowing which part of the state corresponds to which thing. Causal-JEPA restructures the representation around objects, so that masking and prediction operate on entities rather than pixels, and the learned dynamics become modular and intervention-friendly.

How it works

Videoobject slots · object-levelContext encoderf_θTarget encoderf̄_θ · EMAPredictorg_φlatent loss‖ẑ − sg(z̄)‖²z_ctxẑz̄ (sg)EMA copyaction aₜ
Canonical JEPA schematic for Video. The input is split into a visible context and hidden targets (object slot-level, object-level). The context encoder $f_\theta$ embeds what is visible; the target encoder $\bar f_\theta$ (an EMA copy, gradient stopped) embeds the targets; the predictor $g_\phi$ maps context to the target embeddings; training minimises the latent distance. The predictor is action-conditioned: $\hat z_{t+1}=g_\phi(z_t,a_t)$ — this is what turns a representation learner into a world model.

The context encoder produces an object- or entity-centric set of latents — one slot per entity — rather than a flat grid of patch tokens. The self-supervised signal comes from masking at the level of objects: a whole entity's latent is hidden, and the predictor must infer it from the remaining entities and from the action. Training keeps the JEPA core — predict target-encoder latents under stop-gradient with anti-collapse — but the factored structure means the transition $\hat z_{t+1}=g_\phi(z_t,a_t)$ operates on disentangled slots.

Because each slot tracks one entity, the model isolates how that entity (and any intervention on it) propagates through the scene. This yields causal, modular dynamics and makes counterfactual edits — change one slot, predict the consequence — natural to express.

The objective

With slot set $\{z_t^{(i)}\}$ and a masked entity $m$, the loss predicts the masked slot's target latent from context slots and action:

$$\mathcal{L} = \big\lVert g_\phi\big(\{z_t^{(i)}\}_{i\neq m},\, a_t\big) - \operatorname{sg}\big[\bar f_\theta\big(o^{(m)}\big)\big] \big\rVert^2 + \lambda\,\mathcal{R}(Z).$$

The transition acts per slot, $\hat z_{t+1}^{(i)}=g_\phi(z_t,a_t)^{(i)}$, so reasoning over a sparse set of entities replaces dense patch prediction. Anti-collapse keeps the slots informative, and the object structure makes the cost of evaluating an intervention scale with the number of entities, not pixels.

Key results & what's novel

The novelty is object-level latent masking to obtain object-centric world models. Factoring the latent into causal entities supports counterfactual reasoning — silence or alter one slot and predict the rest — and makes planning more efficient, since search and intervention evaluation over a handful of entities is combinatorially cheaper and more interpretable than dense pixel-space rollouts. By aligning masking with the world's causal units, Causal-JEPA moves JEPA world models toward systematic generalisation and explicit intervention semantics rather than mere appearance prediction.

Strengths & limitations

  • + Disentangled entity slots enable clean counterfactual queries.
  • + Object-level reasoning makes planning and intervention search cheaper and more interpretable.
  • + Modular dynamics aid systematic generalisation across scene configurations.
  • − Requires reliable object/slot discovery, which is itself hard in cluttered or deformable scenes.
  • − The number of slots and their binding to entities are design choices that may not match the true causal structure.
  • − Object-centric encoders can struggle with stuff (backgrounds, fluids) that lacks discrete entities.

Connections & references