Motivation
The paper confronts a foundational gap: animals and humans acquire usable world models from modest experience and then plan and reason with them, while machine systems are either brittle pattern matchers or reinforcement learners that need enormous interaction. LeCun argues this gap is architectural, not merely a matter of scale. What is missing is a system that learns predictive models of the world and uses them to imagine the consequences of actions before acting. The position is that intelligent behavior should emerge from a single, end-to-end differentiable design rather than a patchwork of hand-engineered components, and that the central learning problem is self-supervised acquisition of structure from passive observation.
How it works
The blueprint organizes cognition into cooperating modules:
- A configurator that conditions all other modules for the task at hand.
- A perception module estimating the current state of the world from observations.
- A world model that predicts plausible future states and fills in missing information.
- A cost module producing a scalar measuring discomfort or energy.
- An actor that optimizes sequences of actions to minimize predicted future cost.
- A short-term memory storing states, costs, and predictions.
Action selection is inference: the actor searches for an action sequence that the world model predicts will minimize cumulative cost. Multiple world models can be stacked hierarchically so prediction and planning operate at increasing levels of abstraction and longer horizons.
The setup
The proposal is cast in the language of energy-based models. Instead of predicting raw observations, the system learns a scalar energy $E(x,y)$ that measures the incompatibility between a context $x$ and a candidate $y$; compatible configurations get low energy. The decisive contribution is the Joint Embedding Predictive Architecture (JEPA): encode both context and target into representations $s_x = \mathrm{Enc}(x)$ and $s_y = \mathrm{Enc}(y)$, and predict $s_y$ from $s_x$ in latent space, $\hat s_y = \mathrm{Pred}(s_x, z)$. A latent variable $z$ absorbs the information about the target not determined by the context, so the predictor can stay confident while the world remains uncertain — sidestepping the impossibility of predicting every low-level detail of a high-dimensional signal.
Key results & what's novel
This is a position paper, so its contribution is conceptual rather than empirical: it names and frames the JEPA idea and supplies the architecture the whole subsequent family instantiates. The novel claims are that latent prediction should replace generation for world modeling, that self-supervised learning of energy landscapes — non-contrastive where possible — is the right training principle, and that planning reduces to model-predictive optimization over actions in the learned latent space. It also articulates representation collapse as the central failure mode to design against, and proposes hierarchical, multi-timescale world models as the route to long-horizon reasoning.
Strengths & limitations
- + Provides a unified, modular vocabulary that organizes an entire research programme.
- + Latent prediction over generation is now validated empirically by I-JEPA, V-JEPA and successors.
- + Treats uncertainty and planning in a single coherent energy framework.
- − A blueprint, not an implementation: it specifies what to build, not the training recipes that make each module work.
- − Several pieces (the configurator, hierarchical world models, robust planning) remained largely unrealized at publication and are still open problems.
- − Collapse-avoidance, the choice of latent-variable handling, and energy shaping are left as design challenges for later work.