LeWM Push-T: My Experiment vs Paper (arXiv 2603.19312v3)

Legend: ✓ = matches · ≈ = close / likely same · ⚠️ = differs, likely affects results · — = not stated in paper

1. Training parameters

MinePaper
Epochs310⚠️
Image size112×112224×224⚠️
EncoderViT-tiny, patch 14, 12 layers, 3 heads, hidden 192, from scratchViT-tiny, patch 14, 12 layers, 3 heads, hidden 192, from scratch✓
Embedding[CLS] → MLP projector (hidden 2048) + BatchNorm, dim 192[CLS] → 1-layer MLP + BatchNorm, dim 192≈
Predictortransformer, 6 layers, 16 heads, dropout 0.1, AdaLN action conditioning, ~10M paramstransformer (“ViT-S”), 6 layers, 16 heads, 10% dropout, AdaLN, ~10M params✓
History size33✓
Sub-trajectory4 frames + 4 blocks of 5 actions4 frames + 4 blocks of 5 actions✓
Frameskip55✓
Batch size128128✓
SIGReg λ0.090.1 (λ ablation peaks near 0.09)≈
SIGReg projections10241024✓
SIGReg knots17T nodes in [0.2, 4] (count not stated; insensitive)—
Optimizer / LR / WDAdamW / 5e-5 / 1e-3——
Precision / grad clipbf16 / 1.0——
Datasetlibrakevin/lewm-pushtDINO-WM Push-T, 20,000 expert episodes≈ (not verified)
Training seeds1 (seed 3072)3⚠️

2. Eval parameters

MinePaper
Episodes5050 (same 50 for all models)✓
Eval budget50 steps50 steps✓
Goal offset25 steps25 steps✓
CEM samples300300✓
CEM iterations3030 (Push-T)✓
CEM top-k3030✓
CEM initial variance1.01✓
Planning horizon5 (= 25 env steps)5 (= 25 env steps)✓
Receding horizon5 (execute full plan)5 (execute full plan)✓
Eval image size112224⚠️
Eval seed42 (default in config/eval/pusht.yaml)——

3. Eval Result (Push-T success rate)

All 10 seeds together (same epoch-3 checkpoint, 500 episodes in total):

Eval seedSuccess rate
084.0%
174.0%
278.0%
384.0%
476.0%
582.0%
674.0%
792.0%
4282.0%
11380.0%
Mean ± std80.6 ± 5.2

The ± is computed the same way as the paper’s. The sample std is ±5.5.

Differences

  1. Epochs: 3 vs 10. Most likely the main cause of the gap.
  2. Image size: 112 vs 224, for both training and eval.