[{"content":" The LeWorldModel training pipeline. Figure from the LeWorldModel paper.\nLeWorldModel (LeWM) is a Joint-Embedding Predictive Architecture (JEPA) that learns a world model end to end from pixels. Like other JEPA world models, it predicts future states in latent space and plans there, without reconstructing pixels.\nLeWM is the first JEPA that trains stably end to end from pixels with only two loss terms: a prediction loss and a SIGReg loss. SIGReg (Sketched-Isotropic-Gaussian Regularizer) was first proposed in the LeJEPA paper (Balestriero \u0026amp; LeCun, 2025).\nThe encoder of a JEPA maps each observation (a video frame) to a latent embedding. Without a constraint, the encoder can map every frame to the same vector. Then the predicted embedding always equals the target embedding, so the prediction loss is zero, but the embeddings carry no information about the input. Earlier JEPAs prevent this with other methods, such as an EMA target encoder, stop-gradient, a frozen pretrained encoder, or VICReg-style variance and covariance terms. LeWM prevents representation collapse with SIGReg, therefore LeWM\u0026rsquo;s loss has only one tunable hyperparameter, the SIGReg weight λ\\lambdaλ. PLDM, an earlier end-to-end JEPA, has six.\nSIGReg projects the batch of embeddings onto random 1-D directions and computes the Epps–Pulley test statistic for each direction. The Epps–Pulley test calculates the empirical characteristic function of the projected samples. (The characteristic function φ(t)=E[eitX]\\varphi(t) = \\mathbb{E}[e^{itX}]φ(t)=E[eitX] is the Fourier transform of the distribution of XXX.) It then compares this function with the characteristic function of the standard Gaussian, e−t2/2e^{-t^2/2}e−t2/2, by calculating the squared L2 distance between them, weighted by e−t2/2e^{-t^2/2}e−t2/2 over t∈[0,3]t \\in [0, 3]t∈[0,3]. SIGReg takes the average over all directions. The SIGReg loss, multiplied by λ\\lambdaλ, is added to the prediction loss.\nWhy project onto random 1-D directions? Testing whether a high-dimensional distribution is an isotropic Gaussian N(0,I)\\mathcal{N}(0, I)N(0,I) is hard, but the Epps–Pulley test works well in one dimension. By the Cramér–Wold theorem, a distribution is N(0,I)\\mathcal{N}(0, I)N(0,I) if and only if every 1-D projection of it is N(0,1)\\mathcal{N}(0, 1)N(0,1). So it is enough to test 1-D projections. Each training step samples 1024 new random directions, so the encoder cannot hide a \u0026ldquo;bad\u0026rdquo; direction (a direction where the projected embeddings are not N(0,1)\\mathcal{N}(0, 1)N(0,1)) for long. The LeJEPA paper shows that sampling a modest number of new directions at each step is enough to prevent collapse in practice.\nTo make the training pipeline and SIGReg easier to understand, I asked Claude Code to create two interactive visualizations.\nTraining Pipeline Visualization Use the buttons in the visualizer to move from step to step. Each step shows the tensors and their shapes.\nModules with a \u0026ldquo;+\u0026rdquo; sign can be expanded. Click the \u0026ldquo;+\u0026rdquo; to see inside a module, for example, a ViT block or one of the predictor\u0026rsquo;s conditional blocks.\nOpen in a new tab ↗ SIGReg Characteristic Function Visualization This visualizer shows the Epps–Pulley test for one 1-D projection. Each sample xxx is put on the unit circle at angle t⋅xt \\cdot xt⋅x, and the centroid of the points is the empirical characteristic function at ttt. Move the ttt slider, or click Play, to see the centroid follow (or miss) the target e−t2/2e^{-t^2/2}e−t2/2. Then choose another distribution, for example \u0026ldquo;Collapsed\u0026rdquo;, and see how the SIGReg value grows.\nOpen in a new tab ↗ ","permalink":"https://taot.github.io/posts/2026/10/visualizing-leworldmodel/","summary":"How LeWorldModel prevents representation collapse with SIGReg, with interactive visualizations of one training step and of the characteristic-function test.","title":"Visualizing LeWorldModel"}]