TL;DR We show how to recover individual world latents up to signed permutation without reconstruction or a decoder, using DSReg (Dependency-Sparsity Regularization), which can be applied post hoc to any LeJEPA checkpoint at no loss over joint training.
Abstract
Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity. Methods without these anchors, including joint-embedding predictive architectures (JEPAs), identify the latent state only up to a linear transformation, so individual latents remain mixed. We close this gap: individual world latents can be provably recovered with no reconstruction, no decoder, and no labels. The key condition is Structural Diversity: different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike. Building on the linear identifiability that LeJEPA provides, we prove that under Structural Diversity, DSReg (Dependency-Sparsity Regularization) recovers individual world latents up to signed permutation, without reconstruction or a decoder. It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training, and establishes the first fully identifiable JEPA that recovers every world latent. Moreover, as a condition on dependency footprints, Structural Diversity is strictly weaker than all structural conditions of prior identifiable latent variable models. Across synthetic regimes, world model probes, learned visual encoders, and external renderers, DSReg preserves dense prediction while improving individual-latent recovery and downstream use with scales.
The setup
The world. The latent state \(z\in\mathbb{R}^d\) is standard Gaussian with stationary Gaussian dynamics, as in the LeJEPA setting. We observe it only through an unknown nonlinear map \(x=g(z)\).
The learner. LeJEPA predicts the embedding of the next observation from that of the current one while keeping the embedding Gaussian. Klindt, LeCun, and Balestriero (2026) prove that at the optimum of this objective it recovers the latent state only up to a rotation, \(h(z)=Qz\) with \(Q\in O(d)\), so each learned variable may mix several world latents.
DSReg. Among all rotations of the learned variables, DSReg picks a rotation under which the observations have the fewest dependencies on the rotated variables. It minimizes the support criterion, which counts the Jacobian entries that are not almost surely zero:
In practice, DSReg minimizes an \(\ell_1\) relaxation of this count over local Jacobians estimated from the observations and the frozen representation, and trains no decoder. The guarantees below concern the count itself and a thresholded version of it.
Theoretical results
The dependency footprint \(\mathcal{S}_i\) of latent \(z_i\) is the set of observed variables that \(z_i\) affects, and Structural Diversity asks that no two latents share a footprint:
When two latents affect exactly the same observed variables, a rotation within the pair changes nothing the dependency structure can see, so recovery from structure alone needs distinct footprints.
Identifiability without reconstruction
Every minimizer recovers each world latent, up to sign and permutation.
AssumesLeJEPA's output \(h(z)=Qz\) with \(Q\in O(d)\), Structural Diversity, and Functional no-cancellation, a faithfulness condition that rules out cancellations hiding genuine Jacobian entries.
WhyA variable that mixes several latents depends on the union of their footprints, so when footprints are distinct, any mixing raises the total count.
Strictly weaker than prior structural conditions
Structural Diversity follows from each earlier condition and implies none of them.
- Structural Sparsityboth forms
- Non-Inclusion\(\mathcal S_i\not\subseteq\mathcal S_j\)
- Structural Variability\(|\mathcal S_i\triangle\mathcal S_j|\ge 2\)
- Structural Diversity\(\mathcal S_i\ne\mathcal S_j\)
Each formula is required for all \(i\ne j\).
ExampleNested footprints, such as a global latent like illumination whose footprint strictly contains those of local latents, violate Non-Inclusion and both forms of Structural Sparsity. Footprints that differ in a single variable violate all four prior conditions. Both patterns satisfy Structural Diversity.
Stable under estimation error
Small Jacobian errors keep every minimizer within \(\delta\) of a signed permutation.
where \(\delta>0\) and \(P\) ranges over signed permutations. The same condition guarantees that minimizers of the thresholded count exist.
Assumesthe conditions of the identifiability result with \(d\ge2\), and bounded rows of the true Jacobian. The estimation error is measurable and at most \(\varepsilon\) in every row, entries are counted only above a threshold \(\tau>\varepsilon\), and \(\sigma\) and \(\rho^*(\delta)\) are positive constants of the true Jacobian and its support pattern.
The identifiability result is formally verified in Lean 4 with mathlib. The Lean proof contains no sorry and uses only Lean's standard axioms.
Empirical results
Recovery follows Structural Diversity
We train encoders from scratch with LeJEPA on a synthetic benchmark whose footprints are distinct but nested. Their variables stay mixed, and DSReg lifts them to near-ceiling recovery (a). Across footprint regimes (b), recovery succeeds wherever Structural Diversity holds, and with identical footprints DSReg recovers only the span of each pair.
Encoders trained from pixels
On convolutional encoders trained from pixels, DSReg reaches the supervised Procrustes oracle, while PCA, Varimax, and FastICA leave the latents mixed (a). The match holds at every width tested as the representation widens beyond the number of world latents (b).
Each recovered latent responds to one factor
On Gaussian 3DShapes, sweeping one factor with the others fixed moves a single DSReg latent, while in LeJEPA coordinates the response spreads across several latents.
Recovered latents support sparse use
Many modules built on a world model act through a few variables at a time, such as a controller moving one object, and a rotation breaks that interface because one learned variable then moves several physical factors at once. On states \(h=Qz\) with a random orthogonal \(Q\), DSReg improves visual editing, sparse control, rollout prediction, and surprise detection.
Scales to large latent dimension
The local Jacobians factor through each anchor's neighborhood and never need to be materialized, so the full procedure reaches latent dimension 8192 on one 48 GB GPU, with recovery declining gently as the dimension grows.
Code
After pip install -e . in the repository, the snippet below fits DSReg to a frozen representation.
from dsreg import fit_dsreg, whiten
h, mean, W = whiten(h) # frozen representation, shape (n, d)
R, info = fit_dsreg(h, x) # observations x, shape (n, p)
z_tilde = h @ R.T # individual latents, up to sign and permutation
BibTeX
% This entry will be updated with the arXiv identifier.
@misc{dsreg2026,
title = {DSReg: Provably Recovering Individual World Latents without Reconstruction},
author = {Zheng, Yujia and Klindt, David and Balestriero, Randall and Sch{\"o}lkopf, Bernhard},
year = {2026}
}








