Motivation
The cleanest way to build a world model is end-to-end from pixels: learn the encoder, the action-conditioned predictor, and the control-relevant latent jointly. In practice this is notoriously unstable, because nothing stops the representation from collapsing to a constant that trivially satisfies the prediction loss. The standard remedy — a stop-gradient EMA teacher/target encoder — works but adds a second moving network, momentum and schedule hyperparameters, and a recurring source of brittleness. LeWorldModel asks whether the teacher can be removed entirely while still learning stable, plannable action-conditioned dynamics.
How it works
A single context encoder embeds pixels into a latent $z_t$, and a predictor advances it under actions, $\hat z_{t+1}=g_\phi(z_t,a_t)$ — learned directly from pixels, end-to-end. Crucially there is no EMA teacher and no stop-gradient asymmetry: the network that produces context latents also produces the prediction targets, so gradients flow through the whole model.
Collapse is prevented not by a teacher but by SIGReg, a distributional regulariser that constrains the latent toward an isotropic Gaussian via random 1-D projections tested against a standard normal. SIGReg supplies the spreading pressure the teacher normally provides, yielding a stable single-network training signal. At test time the model plans by rolling latents forward under candidate actions and scoring against a goal latent.
The objective
The loss has exactly two terms: action-conditioned latent prediction and the anti-collapse regulariser.
$$\mathcal{L} = \underbrace{\big\lVert g_\phi(z_t, a_t) - z_{t+1} \big\rVert^2}_{\text{action-conditioned prediction}} \;+\; \lambda\,\underbrace{\operatorname{SIGReg}(Z)}_{Z \to \mathcal{N}(0,I)}.$$
Both $z_t$ and the target $z_{t+1}$ come from the same encoder, with no EMA copy. SIGReg drives the embedding distribution toward isotropy, removing the need for a momentum teacher, stop-gradient, or schedule. The result is a compact, single-network recipe for controllable latent dynamics.
Key results & what's novel
The contribution is a genuinely stable, teacher-free recipe for action-conditioned latent world models learned directly from raw pixels. By replacing the EMA target encoder with SIGReg's distributional anti-collapse, LeWorldModel collapses the JEPA-for-control pipeline to a single network and a two-term loss, removing a major source of instability and hyperparameter sensitivity while retaining latent dynamics that remain plannable via model-predictive control. It demonstrates that the principled anti-collapse from LeJEPA extends cleanly from representation learning to action-conditioned world modeling.
Strengths & limitations
- + No EMA teacher, stop-gradient, or schedule to tune; stable end-to-end from pixels.
- + Compact two-term objective with a single anti-collapse hyperparameter.
- + Retains plannable latent dynamics for model-based control.
- − SIGReg's isotropic-Gaussian target rests on assumptions that may not suit every domain.
- − The sketching estimator's variance depends on the number of random projections.
- − A deterministic predictor does not represent uncertainty over multiple futures.