At a glance
ProblemLearning an action-conditioned latent world model end-to-end from pixels is unstable: jointly training encoder, predictor, and action conditioning risks collapse, and the usual EMA teacher fix is fragile.
Key ideaDrop the teacher. Use a two-term objective — action-conditioned latent prediction plus SIGReg anti-collapse — to train one network end-to-end from pixels.
ModalityPixels (+ actions)
Target / maskingNo EMA teacher; both branches share weights. Anti-collapse comes from SIGReg, a distributional isotropy regulariser.
Builds onLeJEPA / SIGReg and the action-conditioned latent world-model recipe.
Used forStable, teacher-free model-based control with plannable latent dynamics.

Motivation

The cleanest way to build a world model is end-to-end from pixels: learn the encoder, the action-conditioned predictor, and the control-relevant latent jointly. In practice this is notoriously unstable, because nothing stops the representation from collapsing to a constant that trivially satisfies the prediction loss. The standard remedy — a stop-gradient EMA teacher/target encoder — works but adds a second moving network, momentum and schedule hyperparameters, and a recurring source of brittleness. LeWorldModel asks whether the teacher can be removed entirely while still learning stable, plannable action-conditioned dynamics.

How it works

Pixelspatchs · blocksContext encoderf_θTarget encodershared · SIGRegPredictorg_φlatent loss‖ẑ − sg(z̄)‖²z_ctxẑz̄ (sg)shared weightsaction aₜ
Canonical JEPA schematic for Pixels. The input is split into a visible context and hidden targets (patch-level, blocks). The context encoder $f_\theta$ embeds what is visible; the target encoder $\bar f_\theta$ (an none copy, gradient stopped) embeds the targets; the predictor $g_\phi$ maps context to the target embeddings; training minimises the latent distance. The predictor is action-conditioned: $\hat z_{t+1}=g_\phi(z_t,a_t)$ — this is what turns a representation learner into a world model.

A single context encoder embeds pixels into a latent $z_t$, and a predictor advances it under actions, $\hat z_{t+1}=g_\phi(z_t,a_t)$ — learned directly from pixels, end-to-end. Crucially there is no EMA teacher and no stop-gradient asymmetry: the network that produces context latents also produces the prediction targets, so gradients flow through the whole model.

Collapse is prevented not by a teacher but by SIGReg, a distributional regulariser that constrains the latent toward an isotropic Gaussian via random 1-D projections tested against a standard normal. SIGReg supplies the spreading pressure the teacher normally provides, yielding a stable single-network training signal. At test time the model plans by rolling latents forward under candidate actions and scoring against a goal latent.

The objective

The loss has exactly two terms: action-conditioned latent prediction and the anti-collapse regulariser.

$$\mathcal{L} = \underbrace{\big\lVert g_\phi(z_t, a_t) - z_{t+1} \big\rVert^2}_{\text{action-conditioned prediction}} \;+\; \lambda\,\underbrace{\operatorname{SIGReg}(Z)}_{Z \to \mathcal{N}(0,I)}.$$

Both $z_t$ and the target $z_{t+1}$ come from the same encoder, with no EMA copy. SIGReg drives the embedding distribution toward isotropy, removing the need for a momentum teacher, stop-gradient, or schedule. The result is a compact, single-network recipe for controllable latent dynamics.

Key results & what's novel

The contribution is a genuinely stable, teacher-free recipe for action-conditioned latent world models learned directly from raw pixels. By replacing the EMA target encoder with SIGReg's distributional anti-collapse, LeWorldModel collapses the JEPA-for-control pipeline to a single network and a two-term loss, removing a major source of instability and hyperparameter sensitivity while retaining latent dynamics that remain plannable via model-predictive control. It demonstrates that the principled anti-collapse from LeJEPA extends cleanly from representation learning to action-conditioned world modeling.

Strengths & limitations

  • + No EMA teacher, stop-gradient, or schedule to tune; stable end-to-end from pixels.
  • + Compact two-term objective with a single anti-collapse hyperparameter.
  • + Retains plannable latent dynamics for model-based control.
  • − SIGReg's isotropic-Gaussian target rests on assumptions that may not suit every domain.
  • − The sketching estimator's variance depends on the number of random projections.
  • − A deterministic predictor does not represent uncertainty over multiple futures.

Connections & references