At a glance
ProblemJEPAs learn useful latents, but it is unclear when a JEPA recovers the true generative world variables rather than an arbitrary tangle.
Key ideaUnder stated assumptions, prove a JEPA linearly recovers latent world variables up to rotation — read this as design guidance, not a guarantee.
ModalityTheory (domain-agnostic)
Target / maskingStandard JEPA prediction with LeJEPA-style Gaussian regularisation; no EMA teacher required for the result.
Builds onLeJEPA (isotropic-Gaussian embeddings, SIGReg) and identifiability theory.
Used forUnderstanding which inductive biases make JEPA latents plannable; connects to optimal latent-space planning.

Motivation

JEPAs empirically produce latents that support recognition and even planning, but empirical usefulness does not tell us what the latent represents. Does a JEPA recover the underlying generative variables of the world, or some arbitrarily entangled, rotationally scrambled function of them? The distinction matters for control: faithful, structured latent transitions are what make planning meaningful. This work asks, for the LeJEPA setting, under exactly which conditions the learned representation is identifiable with the true latent world variables — turning a vague intuition into a precise, checkable claim.

How it works

Any modalitytokens · blocksContext encoderf_θTarget encodershared · SIGRegPredictorg_φlatent loss‖ẑ − sg(z̄)‖²z_ctxẑz̄ (sg)shared weightsaction aₜ
Canonical JEPA schematic for Any modality. The input is split into a visible context and hidden targets (token-level, blocks). The context encoder $f_\theta$ embeds what is visible; the target encoder $\bar f_\theta$ (an none copy, gradient stopped) embeds the targets; the predictor $g_\phi$ maps context to the target embeddings; training minimises the latent distance. The predictor is action-conditioned: $\hat z_{t+1}=g_\phi(z_t,a_t)$ — this is what turns a representation learner into a world model.

The analysis sets up a generative model of the world and a JEPA — context encoder plus latent predictor — trained with LeJEPA-style Gaussian regularisation of the embedding. It then asks what the optimum of that objective implies about the relationship between the learned latent and the true latent. The key ingredients are three assumptions: the true latents are independent and Gaussian, transitions are stationary with additive noise, and the embedding is successfully regularised toward isotropic Gaussianity (LeJEPA's anti-collapse, so no EMA teacher is needed). Under these, the prediction objective constrains the encoder enough that the learned latent must be a structured transform of the truth.

The result

The theorem states that, under the stated assumptions, a JEPA linearly recovers the latent world variables up to a rotation: the learned latent space is an orthogonal transform $R$ of the true one,

$$\hat z = R\, z_{\text{true}}, \qquad R^\top R = I.$$

Identifiability up to rotation means no information about the latent state is lost or distorted beyond a rigid reorientation. Consequently the latent transitions are faithful — an additive-noise transition stays additive under $R$ — and the geometry supports optimal latent-space planning, with $\hat z_{t+1}=g_\phi(z_t,a_t)$ tracking the true dynamics up to that fixed rotation.

Why it matters

The result gives a rigorous link between JEPA's self-supervised objective and world-model identifiability, explaining why Gaussian-regularised latent prediction can yield plannable, structured representations rather than collapsed or rotationally arbitrary ones. It elevates isotropic Gaussian regularisation from a mere collapse fix to a principled choice with provable consequences, and it tells practitioners which inductive biases — independent factors, stationary additive transitions, isotropy — to aim for if they want latents whose geometry is trustworthy for planning.

Assumptions & caveats

This is design guidance, not a guarantee. The assumptions are strong and frequently violated in practice. Real generative factors are correlated and nonlinear, not independent Gaussians; many transitions are non-stationary, multiplicative, or state-dependent rather than additive-noise; and Gaussian regularisation is only approximately achieved. When these conditions break, recovery degrades from exact-up-to-rotation to merely approximate, and identifiability can fail outright.

  • + Precise conditions connecting the JEPA objective to plannable latents.
  • − Independence, Gaussianity, and additive-noise stationarity rarely hold exactly.
  • − Up-to-rotation identifiability is a best case; real systems land short of it.

Connections & references