DSReg: Provably Recovering Individual World Latents without Reconstruction

Yujia Zheng1 David Klindt2 Randall Balestriero3 Bernhard Schölkopf4,5

1University of Illinois Urbana-Champaign 2Cold Spring Harbor Laboratory 3Brown University 4Max Planck Institute for Intelligent Systems 5ELLIS Institute Tübingen

Animation based on Figure 1. Four world latents, a sun, a moon, a thermometer and a cloud with wind, generate one scene. The learned variables start as two-latent mixtures drawn as split circles, above a grid showing which observed variables each one depends on, with 14 dependencies in total. DSReg rotates them until each circle holds a single latent, some with a minus sign, and the count falls to 10.
A LeJEPA world model can learn variables that mix hidden factors (split circles), and DSReg rotates them until they affect the fewest observed variables (dots, 14 down to 10). When no two factors affect exactly the same observed variables, this separates the factors, one per learned variable, possibly reordered and flipped in sign.

TL;DR We show how to recover individual world latents up to signed permutation without reconstruction or a decoder, using DSReg (Dependency-Sparsity Regularization), which can be applied post hoc to any LeJEPA checkpoint at no loss over joint training.

Abstract

Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity. Methods without these anchors, including joint-embedding predictive architectures (JEPAs), identify the latent state only up to a linear transformation, so individual latents remain mixed. We close this gap: individual world latents can be provably recovered with no reconstruction, no decoder, and no labels. The key condition is Structural Diversity: different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike. Building on the linear identifiability that LeJEPA provides, we prove that under Structural Diversity, DSReg (Dependency-Sparsity Regularization) recovers individual world latents up to signed permutation, without reconstruction or a decoder. It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training, and establishes the first fully identifiable JEPA that recovers every world latent. Moreover, as a condition on dependency footprints, Structural Diversity is strictly weaker than all structural conditions of prior identifiable latent variable models. Across synthetic regimes, world model probes, learned visual encoders, and external renderers, DSReg preserves dense prediction while improving individual-latent recovery and downstream use with scales.

The setup

Four panels. World latents z1 to z4 are drawn as a sun, a moon, a thermometer and a cloud with wind. The observation x = g(z) is one scene that combines them. LeJEPA latents are mixtures of two factors, such as a z1 + b z2. DSReg latents are the four individual factors in a permuted order.
Figure 1DSReg turns a mixed latent state into individual latents. Four world latents \(z\) (sun, moon, temperature, wind) generate the observed scene \(x=g(z)\). The variables LeJEPA learns may each mix several latents, and after DSReg each variable is a single latent, possibly reordered and flipped in sign.

The world. The latent state \(z\in\mathbb{R}^d\) is standard Gaussian with stationary Gaussian dynamics, as in the LeJEPA setting. We observe it only through an unknown nonlinear map \(x=g(z)\).

The learner. LeJEPA predicts the embedding of the next observation from that of the current one while keeping the embedding Gaussian. Klindt, LeCun, and Balestriero (2026) prove that at the optimum of this objective it recovers the latent state only up to a rotation, \(h(z)=Qz\) with \(Q\in O(d)\), so each learned variable may mix several world latents.

DSReg. Among all rotations of the learned variables, DSReg picks a rotation under which the observations have the fewest dependencies on the rotated variables. It minimizes the support criterion, which counts the Jacobian entries that are not almost surely zero:

\[ \min_{R\in O(d)} \Big\|\frac{\partial x}{\partial\tilde z}\Big\|_{0,\mu}, \qquad \tilde z=Rh(z). \]
\[ \begin{gathered} \min_{R\in O(d)} \Big\|\frac{\partial x}{\partial\tilde z}\Big\|_{0,\mu},\\ \tilde z=Rh(z). \end{gathered} \]

In practice, DSReg minimizes an \(\ell_1\) relaxation of this count over local Jacobians estimated from the observations and the frozen representation, and trains no decoder. The guarantees below concern the count itself and a thresholded version of it.

Theoretical results

The dependency footprint \(\mathcal{S}_i\) of latent \(z_i\) is the set of observed variables that \(z_i\) affects, and Structural Diversity asks that no two latents share a footprint:

\[\mathcal S_i \ne \mathcal S_j \quad \text{for all } i \ne j.\]

When two latents affect exactly the same observed variables, a rotation within the pair changes nothing the dependency structure can see, so recovery from structure alone needs distinct footprints.

Top: a graph with arrows from latents z1, z2, z3 to observed variables x1 to x4. Bottom: the matching support matrix, with a dot where a latent affects an observed variable. Each latent column has a different pattern of dots.
Figure 2Dependency footprints. Each column of \(\partial x/\partial z\) marks which observed variables one latent affects.

Identifiability without reconstruction

Every minimizer recovers each world latent, up to sign and permutation.

\[ R \in \arg\min_{R\in O(d)} \Big\|\frac{\partial x}{\partial \tilde z}\Big\|_{0,\mu} \;\Longrightarrow\; RQ \text{ is a signed permutation, } \tilde z_i = s_i\, z_{\pi(i)} \text{ a.s.} \]
\[ \begin{gathered} R \in \arg\min_{R\in O(d)} \Big\|\frac{\partial x}{\partial \tilde z}\Big\|_{0,\mu} \\ \Longrightarrow\; RQ \text{ is a signed permutation,} \\ \tilde z_i = s_i\, z_{\pi(i)} \text{ a.s.} \end{gathered} \]

AssumesLeJEPA's output \(h(z)=Qz\) with \(Q\in O(d)\), Structural Diversity, and Functional no-cancellation, a faithfulness condition that rules out cancellations hiding genuine Jacobian entries.

WhyA variable that mixes several latents depends on the union of their footprints, so when footprints are distinct, any mixing raises the total count.

Strictly weaker than prior structural conditions

Structural Diversity follows from each earlier condition and implies none of them.

  1. Structural Sparsityboth forms
  2. Non-Inclusion\(\mathcal S_i\not\subseteq\mathcal S_j\)
  3. Structural Variability\(|\mathcal S_i\triangle\mathcal S_j|\ge 2\)
  4. Structural Diversity\(\mathcal S_i\ne\mathcal S_j\)

Each formula is required for all \(i\ne j\).

ExampleNested footprints, such as a global latent like illumination whose footprint strictly contains those of local latents, violate Non-Inclusion and both forms of Structural Sparsity. Footprints that differ in a single variable violate all four prior conditions. Both patterns satisfy Structural Diversity.

Stable under estimation error

Small Jacobian errors keep every minimizer within \(\delta\) of a signed permutation.

\[ \tau+\varepsilon<\sigma\,\rho^*(\delta) \;\Longrightarrow\; \min_{P}\,\|\widehat R Q-P\|_F<\delta , \]
\[ \begin{gathered} \tau+\varepsilon<\sigma\,\rho^*(\delta) \\ \Longrightarrow\; \min_{P}\,\|\widehat R Q-P\|_F<\delta , \end{gathered} \]

where \(\delta>0\) and \(P\) ranges over signed permutations. The same condition guarantees that minimizers of the thresholded count exist.

Assumesthe conditions of the identifiability result with \(d\ge2\), and bounded rows of the true Jacobian. The estimation error is measurable and at most \(\varepsilon\) in every row, entries are counted only above a threshold \(\tau>\varepsilon\), and \(\sigma\) and \(\rho^*(\delta)\) are positive constants of the true Jacobian and its support pattern.

The identifiability result is formally verified in Lean 4 with mathlib. The Lean proof contains no sorry and uses only Lean's standard axioms.

Empirical results

Recovery follows Structural Diversity

We train encoders from scratch with LeJEPA on a synthetic benchmark whose footprints are distinct but nested. Their variables stay mixed, and DSReg lifts them to near-ceiling recovery (a). Across footprint regimes (b), recovery succeeds wherever Structural Diversity holds, and with identical footprints DSReg recovers only the span of each pair.

Line plot of latent MCC against latent dimension N from 4 to 14. DSReg (solid teal) stays near the ceiling at every N, and LeJEPA (dashed gray) falls as N grows.

(a) Learned encoders

Four bar charts of latent MCC for LeJEPA and DSReg under diverse, nested, minimal-difference and identical footprints. Labels above the charts say which prior conditions fail, and that Structural Diversity holds in the first three regimes and fails in the last. DSReg is near the dashed oracle line in the first three. With identical footprints it stays well below the line, and a label reports that it recovers the span of each pair.

(b) Footprint regimes (\(N=16\))

Figure 3DSReg recovers individual latents wherever Structural Diversity holds, and only there. (a) DSReg (solid) is near ceiling at every \(N\), twenty runs. LeJEPA (dashed) stays mixed. (b) On analytic orbits \(h=Qz\), recovery succeeds wherever the condition holds and falls back to the shared pair subspace where it fails.

Encoders trained from pixels

On convolutional encoders trained from pixels, DSReg reaches the supervised Procrustes oracle, while PCA, Varimax, and FastICA leave the latents mixed (a). The match holds at every width tested as the representation widens beyond the number of world latents (b).

Two bar charts of latent MCC, visual factors and object scenes, for LeJEPA, PCA, Varimax, FastICA and DSReg, with a dashed line for the supervised oracle. Only DSReg reaches the oracle line.

(a) Rotation baselines on learned encoders

Line plot of latent MCC against the estimated width, dim z tilde, at 8, 16 and 32. DSReg stays high at every width and LeJEPA stays lower.

(b) Overparameterized: \(\dim\tilde z\) grows

Figure 4DSReg matches the supervised Procrustes oracle on encoders trained from pixels. (a) Latent-only rotations leave the variables mixed. DSReg closes the gap to the label-fitted oracle (fifteen seeds). (b) With \(\dim z=8\) fixed, the match holds within \(0.001\) as the estimate widens to \(\dim\tilde z=32\) (five seeds).

Each recovered latent responds to one factor

On Gaussian 3DShapes, sweeping one factor with the others fixed moves a single DSReg latent, while in LeJEPA coordinates the response spreads across several latents.

Six rows, one per 3DShapes factor: floor hue, wall hue, object hue, scale, shape and orientation. Each row shows rendered frames of a sweep of that factor, then two line plots of every learned latent's response against the sweep, one for LeJEPA and one for DSReg. In the DSReg plots a single teal curve follows the factor while the gray curves of the other latents stay flat. In the LeJEPA plots the best-matching latent tracks the factor less closely and several gray curves also move.
Figure 5The rotation turns entangled responses into one selective response per factor. One run of Gaussian 3DShapes. Each row sweeps one factor with all others fixed (rendered frames on the left); the right panels show every learned latent's standardized response in LeJEPA and in DSReg coordinates, on a shared scale per row. The flat gray curves carry the no-mixing claim across all six factors.

Recovered latents support sparse use

Many modules built on a world model act through a few variables at a time, such as a controller moving one object, and a rotation breaks that interface because one learned variable then moves several physical factors at once. On states \(h=Qz\) with a random orthogonal \(Q\), DSReg improves visual editing, sparse control, rollout prediction, and surprise detection.

Six bar charts comparing LeJEPA in gray and DSReg in teal: editing success on TwoRoom, FetchSlide and PushT, and on PushT the MPC success, rollout R squared and surprise AUROC. DSReg is higher in every panel.
Figure 6DSReg improves all six sparse-use probes. DSReg in teal, LeJEPA in gray. Tasks are (1) editing (top) and (2) control, prediction, and monitoring (bottom).

Scales to large latent dimension

The local Jacobians factor through each anchor's neighborhood and never need to be materialized, so the full procedure reaches latent dimension 8192 on one 48 GB GPU, with recovery declining gently as the dimension grows.

Four panels. (a) Latent MCC against latent dimension d from 128 to 8k for dense and sampled evaluation. Both decline gently, and open markers continue the sampled curve at 4k and 8k. (b) MCC against the neighborhood ratio k/d for d = 256, 1024 and 4096. Recovery fails at k/d of 1 for every d, fails at every smaller ratio for the two larger d, and is high from 1.25 upward. (c) At d = 2048, MCC is flat from 16 to 256 anchors. (d) Step time on a log scale. The sampled curve is nearly flat in d while the dense curve rises steeply and crosses it between 512 and 1k.
Figure 7Scaling the full estimated procedure. (a) Recovery declines gently with \(d\); open markers are the single-GPU frontier, where anchor count shrinks to fit memory. (b) Recovery against the neighborhood ratio \(k/d\) at three scales; dashed lines mark the analytic score of an unrotated mixing, \(\sqrt{2\ln d/d}\). (c) Sixteen anchors already match \(256\) at \(d=2048\). (d) The sampled step cost is nearly flat in \(d\) while the dense (materialized) cost grows two orders of magnitude, with the curves crossing between \(d=512\) and \(d=1024\).

Code

After pip install -e . in the repository, the snippet below fits DSReg to a frozen representation.

from dsreg import fit_dsreg, whiten
h, mean, W = whiten(h)          # frozen representation, shape (n, d)
R, info = fit_dsreg(h, x)       # observations x, shape (n, p)
z_tilde = h @ R.T               # individual latents, up to sign and permutation

BibTeX

dsreg2026
% This entry will be updated with the arXiv identifier.
@misc{dsreg2026,
  title  = {DSReg: Provably Recovering Individual World Latents without Reconstruction},
  author = {Zheng, Yujia and Klindt, David and Balestriero, Randall and Sch{\"o}lkopf, Bernhard},
  year   = {2026}
}