At a glance
ProblemBehavior-cloning policies regress raw actions without any predictive model of how observations evolve, making them data-hungry and brittle.
Key ideaTrain a JEPA whose predictor jointly forecasts latent representations of future actions and future observations, injecting world-model structure into action learning.
ModalityObservation + action sequences (imitation / policy learning)
Target / maskingTemporal masking; predict target-encoder latents of future actions and observations under stop-gradient.
Builds onThe JEPA latent-prediction template (I-JEPA / V-JEPA) and action-chunking imitation learning.
Used forSample-efficient policy learning and as a controllable latent world model for planning.

Motivation

Imitation learning typically maps observations straight to actions and regresses them, with no internal account of consequences. Such policies need many demonstrations and break under distribution shift, because nothing forces the model to understand how the world responds to what it does. ACT-JEPA argues that a policy is most data-efficient when action learning is coupled to a predictive model of observations: an agent that can anticipate both its next move and the state that follows has learned intent and consequence together. The aim is to import world-model structure into action representation learning while keeping a single self-supervised, JEPA-style objective rather than bolting on a separate dynamics model.

How it works

Obs+Actiontokens · temporalContext encoderf_θTarget encoderf̄_θ · EMAPredictorg_φlatent loss‖ẑ − sg(z̄)‖²z_ctxẑz̄ (sg)EMA copyaction aₜlocal loss (e.g. MLM)
Canonical JEPA schematic for Obs+Action. The input is split into a visible context and hidden targets (token-level, temporal). The context encoder $f_\theta$ embeds what is visible; the target encoder $\bar f_\theta$ (an EMA copy, gradient stopped) embeds the targets; the predictor $g_\phi$ maps context to the target embeddings; training minimises the latent distance. The predictor is action-conditioned: $\hat z_{t+1}=g_\phi(z_t,a_t)$ — this is what turns a representation learner into a world model. A local/generative loss runs alongside latent prediction (hybrid objective).

A context encoder embeds the agent's history of observations and actions into latents. A predictor is then trained to jointly predict the latent representations of future actions and future observations, rather than reconstructing raw control signals or pixels. Concretely, given context latents and a horizon, it produces $\hat z^{\text{act}}_{t+1:t+H}$ and $\hat z^{\text{obs}}_{t+1:t+H}$, matched to target-encoder embeddings of the true future actions and observations.

Masking over time provides the self-supervised signal: hidden future steps must be inferred from visible context, so the model learns abstract action chunks together with the latent dynamics they induce. Because the predictor already advances observation latents under actions, $\hat z_{t+1}=g_\phi(z_t,a_t)$, the trained policy doubles as a controllable world model usable for planning.

The objective

The loss is a sum of two latent-prediction terms, for the action and observation streams, against stop-gradient targets:

$$\mathcal{L} = \big\lVert g_\phi^{\text{act}}(z_{\text{ctx}}) - \operatorname{sg}[\bar f_\theta(a_{t+1:t+H})] \big\rVert^2 + \big\lVert g_\phi^{\text{obs}}(z_{\text{ctx}}) - \operatorname{sg}[\bar f_\theta(o_{t+1:t+H})] \big\rVert^2.$$

An anti-collapse mechanism keeps both streams from degenerating to constants. Predicting in latent space — rather than regressing raw actions — lets the model capture the abstract, multi-step structure of behavior while remaining a single JEPA objective.

Key results & what's novel

The novelty is making the predictor the locus of policy learning: by jointly forecasting action and observation latents, ACT-JEPA learns representations that encode both what to do and what will happen, improving policy quality and sample efficiency over pure behavior cloning. Because the same model predicts future observation latents under actions, it is implicitly a world model — the learned policy can be used for model-based planning, not just reactive control. The work shows that world-model structure and policy learning need not be separate systems; one latent-prediction objective can serve both.

Strengths & limitations

  • + Couples intent and consequence in a single self-supervised objective.
  • + More sample-efficient and robust than raw-action behavior cloning.
  • + The trained model is reusable as a controllable world model for planning.
  • − Still relies on demonstration data to ground the action stream.
  • − The deterministic predictor does not represent uncertainty over multiple plausible futures.
  • − Jointly balancing the action and observation losses adds a tuning surface.

Connections & references

Builds onV-JEPAI-JEPA