A pose-conditioned, autoregressive Diffusion Transformer with frame-causal attention and spatial memory, in the style of World Labs' RTFM, built at nano scale and run under a hard budget of 20 GPU-hours on Kaggle's 2× Tesla T4 for the entire project. The deliverable is the measurement, not the model: which frontier training methods (REPA, Muon, muP, diffusion forcing) actually pay off at ~15–40M parameters, and by what factor, when every GPU-second is ticketed and ledgered. This is a speedrun / ablation project in the modded-nanogpt tradition, not a scaling project.
Status (2026-09-18). Milestones M0–M2 closed, the M3 ablation (REPA × optimiser, 2 seeds, paired, equal wall-clock) is built, pre-flighted and awaiting its budget ticket. 1.27 of 20 GPU-hours spent. No headline effect size yet; what exists is the measurement apparatus, the throughput and learning-rate facts it produced, and a record of every run that failed on the way.
The question is deliberately narrow. World-model papers ship with a bundle of training choices, each individually plausible, none individually costed. At 20 GPU-hours the bundle cannot be afforded, so the choices have to be measured one at a time, on paired seeds, at equal compute, or the result is anecdote.
Hard constraints, all of which shaped the code:
- Hardware: Kaggle, 2× Tesla T4 (sm75): no bf16, no FlashAttention-2. fp16 AMP with GradScaler, a NaN watchdog, Muon's Newton–Schulz in fp32.
- Sessions: 12 h hard cutoff; checkpoint/resume must survive SIGTERM.
- Budget: 10 wall-clock hours on 2 GPUs = 20 GPU-hours, pre-allocated per
milestone (
budget/allocation.yaml). Every training run needs a ticket (duration, purpose, expected result, abort criterion) approved before launch and appends tobudget/ledger.jsonl; a run that exceeds its ticket is aborted by the trainer itself. Inference (latent precompute, evaluation) is charged to the separate Kaggle weekly quota, not to the training budget. - Data: procedural indoor scenes with exact camera poses (5–15 primitives per room, deterministic per seed, looping trajectories so that a revisit metric is possible), rendered with moderngl or a numpy/numba fallback. No real datasets: pose noise would contaminate the ablations.
flowchart LR
S[Procedural scenes<br/>exact poses, looping paths] --> R[Renderer<br/>moderngl / numba]
R --> AE[Frozen DC-AE f64<br/>16 tokens per frame]
R --> D[DINOv2-small<br/>REPA targets]
AE --> L[(Latent cache<br/>Kaggle dataset)]
D --> L
L --> T[Causal DiT<br/>Plücker raymap, adaLN-single,<br/>RMSNorm, QK-norm, SwiGLU]
T --> O[Rectified flow +<br/>diffusion forcing]
O --> E[Eval: revisit-PSNR, drift,<br/>noise floor from M0 ceiling]
B[Budget ledger + tickets] -. gates every run .-> T
Model (src/model/). A causal DiT: tokens of frame t attend to all tokens of
frames ≤ t, bidirectionally within a frame. Camera pose enters as a Plücker raymap;
conditioning via adaLN-single with zero-initialised modulation so the conditioning
path starts as an exact no-op (an fp16 mitigation, see the module docstring). Every
frame carries its own rectified-flow timestep (diffusion forcing). KV-cache and
frustum-overlap retrieval are implemented for rollouts. Parameter count is
independent of tokens-per-frame and context length, which is what let the
autoencoder decision be taken separately from the model-size gates. Presets: 5M,
15M (dim 320, depth 11), 40M.
Training (src/train/). DDP, fp16 AMP, NaN watchdog, EMA in fp32, muP
(mup.py, source-equivalent to the reference: MuReadout forward multiplier,
1/d_head attention, fan-in-derived LR groups over the actual proxy/target shapes),
Muon (muon.py), REPA alignment head (repa.py, applied inside the model's forward
so its gradient is all-reduced under DDP), and a budget guard that refuses tickets
exceeding the milestone or total allocation.
Evaluation (src/eval/). Autoencoder ceiling, sampler, PSNR/LPIPS, and the
revisit/persistence metric the project was designed to add to the Target-Bench
harness (TUM-AVS, arXiv 2511.17792). Every model number is reported against the M0
ceiling and budget-normalised ("PSNR x after y GPU-minutes"), never bare.
Process. Six narrow-scope agents (data, model, training infrastructure, T4
profiler, evaluation, and a report-only skeptic with veto power at every gate) work
from CLAUDE.md, which doubles as the decision log; the user approves every ticket.
| candidate | tokens/frame | PSNR mean | PSNR min | LPIPS mean |
|---|---|---|---|---|
| DC-AE f32 (SANA) @ 128 px | 16 | 40.68 dB | 29.38 dB | 0.0079 |
| DC-AE f64 (ImageNet) @ 256 px | 16 | 46.01 dB | 35.25 dB | 0.0043 |
| DC-AE f32 (SANA) @ 256 px | 64 | 44.29 dB | 34.78 dB | 0.0045 |
The reconstruction-trained autoencoder beat the diffusion-backbone-trained one on pure fidelity at a quarter of the tokens. 46.0 dB / 0.0043 is the hard ceiling for every later number.
Steps that fit each milestone, compiled, batch 8:
| tokens/frame | M3 per arm (1.25 GPU-h, 15M) | M4 (8 GPU-h, 40M) |
|---|---|---|
| 16 | 151 873 | 475 091 |
| 64 | 29 479 | 97 889 |
| 256 | 2 639 | OOM |
Three facts that overturned the analytical plan: torch.compile is mandatory
(1.2–2.6× speed-up, up to 40 % less VRAM, compiled on sm75 in every configuration);
256 tokens/frame is out of reach, which ruled out SD-VAE at 256 px; and MFU is 4–14 %
everywhere, so these models are launch-overhead-bound and larger batches are worth
more than they would be in a compute-bound regime. fp16 was stable: zero non-finite
losses, zero GradScaler skips.
Eight Kaggle attempts were needed to get sampler and evaluation running end to end,
each on a distinct real bug (AE scaling factor, dataset mount timing, torch.load
default, torch.compile renaming checkpoint keys, context length read from the wrong
axis, a GradScaler growth interval that produced an inf gradient at step ≈ 6 000, a
fit() that never checkpointed on normal completion). The ledger keeps all of them.
The 5M model then trained to 8 000 steps with zero NaN events, but reconstruction sat
at PSNR 13.1 dB / LPIPS 0.61 against the 46.0 dB ceiling, flat across four
checkpoints spanning a 2× step range. A budget-free local diagnostic
(scripts/diagnose_m1_flat_psnr.py) showed why, and that it was not the failure the
gate was designed to catch:
- perturbing one frame's pose changes that frame's prediction by ~4× its magnitude and leaks exactly nothing into earlier frames: conditioning is live, the causal mask is exact;
- one-step denoising near the data manifold is excellent (MSE 0.01 % of a random baseline) and degrades toward pure noise (27 % at t = 0.98);
- sampling from noise gives the same latent MSE at 50, 200 and 1 000 Euler steps (0.978 / 0.980 / 0.981), so integration accuracy is not the cause.
The learned velocity field is good locally and poor globally, because M1's dataset was eight repeated windows of one 128-frame trajectory. Conclusion recorded: conditioning works; do not spend more M1 budget chasing PSNR on that data. 0.78 of 1.0 GPU-hours.
The first muP draft was stopped before touching a GPU because it was not
source-equivalent (depth and head count changed with width, fused QKV omitted from LR
scaling, readout scaled on the wrong side). The corrected same-depth pair, compiled,
zero divergence (results/m2_profile.json):
| preset | dim | params | step time | steps / GPU-h |
|---|---|---|---|---|
| m2_proxy_5m | 192 | 5.33M | 15.8 ms | 228 571 |
| 15m (heads = 8) | 320 | 14.80M | 27.7 ms | 130 105 |
Ten-arm sweep, 5 log-spaced LRs × 2 widths, 3 000 steps, single seed, 0.40 GPU-h
(results/m2_lr_sweep.json): best LR 3e-3 at dim 192, 1e-2 at dim 320, a ratio of
3.3 against muP's prediction of ≈ 1. But the top two candidates of each width differ
by under 1.5 % in loss and are adjacent points on a 3.16×-spaced grid; a single-seed,
five-point sweep cannot separate "transfer works, noise picked the neighbour" from
"transfer is off by a few ×". Both lr = 1e-4 arms diverged around step 600,
symmetrically across widths, pointing at a warm-up instability rather than muP. The
gate was treated as passed by user decision, with the caveat carried forward in
the decision log: revisit before trusting width transfer if M3/M4 behave in an
LR-sensitive way. 1.17 of 3.0 GPU-hours.
The M3 matrix is REPA on/off × Muon vs. AdamW at 15M, two seeds, paired on seed, data
order and init, ranked on the flow loss only (a CPU smoke had reported a "+0.49 REPA
effect" that was purely REPA's own auxiliary term), with equal wall-clock per arm
and a per-arm cosine horizon derived from measured throughput, because an arm cut off
long before its total_steps trains at near-peak LR throughout (that exact mistake
diverged an M1 run). Rectified flow vs. DDPM was dropped as the axis with the
strongest prior and therefore the least information per GPU-hour.
Measured on a single T4, compiled, 0.10 GPU-h:
| arm | steps/s | relative |
|---|---|---|
| no REPA, AdamW | 17.63 | 1.00 |
| REPA, AdamW | 16.20 | 0.92 |
| no REPA, Muon | 9.79 | 0.56 |
| REPA, Muon | 8.96 | 0.51 |
Muon costs ~44 % of throughput on T4 (fp32 Newton–Schulz on every 2-D matrix every
step); REPA's head costs ~8 %. Under the equal-wall-clock rule the Muon arms get about
55 % of the AdamW arms' steps. That is the price the ablation charges; if Muon still
wins, it wins per GPU-hour, the only unit this project reports in. A 600-step Muon LR
probe bracketed the reference default (0.005 → 0.503, 0.02 → 0.460, 0.05 → 0.469
with one NaN-skipped step; AdamW at M2's 3e-4 → 0.568), so muon_lr = 0.02 is set
for M3-B, removing the "tuned AdamW vs. default Muon" bias that the decision log had
flagged. The 19 % loss lead of Muon in that probe is at equal steps, before its
throughput cost; M3-B is what settles it.
| milestone | allocated | spent | state |
|---|---|---|---|
| M1 pipeline / overfit | 1.0 | 0.78 | closed by diagnosis |
| M2 muP transfer | 3.0 | 1.17 | closed by decision |
| M3 ablations | 5.0 | 0.10 | pre-flight done, M3-B ticket pending |
| M4 main run (40M) | 8.0 | 0 | |
| M5 rollout fine-tune | 2.0 | 0 | |
| reserve | 1.0 | 0 | |
| total | 20.0 | 1.27 |
- There is no world-model result yet: no revisit-PSNR, no drift curve, no M3 effect size. The repository currently documents the apparatus and its calibration.
- The M2 gate is a practical call under a real ambiguity, not a statistical pass.
- Two seeds per arm is the minimum the work rules allow; effects inside seed spread
will be reported as
consistent_sign: false, not as a mean. - Procedural scenes with flat-shaded primitives are far from any real distribution; this is by design (exact poses, no pose noise) and limits external validity.
- Ledger entries record
git_sha: unknownfor early runs; provenance for those is by commit message and date.
pip install -e . # torch >= 2.x, numpy; pyyaml for scripts/budget.py
pytest -q # 187 tests: shapes, causality, determinism, resume, budget guards, muP, Muon, REPA
python scripts/budget.py status
python scripts/train.py --config configs/smoke_cpu.yaml # CPU smoke of the full train loopKaggle runs use the notebooks in kaggle/ with the source bundled as a dataset; every
run config is a versioned YAML in configs/. The committed M2 and M3 configs carry
ticket_hours: null on purpose and a test enforces it: a concrete ticket is written
only into the Kaggle-side copy after approval, and the runners refuse to start on
unmeasured steps_per_arm.
| Path | Content |
|---|---|
CLAUDE.md |
project brief, work rules and the full decision log |
budget/ |
milestone allocation and the append-only run ledger |
configs/ |
one YAML per run (smoke, M1, M2 sweep, M3 profile, M3 ablation) |
src/data/ |
procedural scenes, trajectories, renderer, DINOv2 features |
src/model/ |
causal DiT, blocks, Plücker raymap, diffusion forcing, KV-cache, retrieval |
src/train/ |
trainer, DDP, AMP, checkpoint/resume, EMA, muP, Muon, REPA, NaN watchdog, budget |
src/eval/ |
AE ceiling, sampler, metrics, test-frame protocol |
scripts/ |
training entry point, budget CLI, precompute, per-milestone runners, M1 diagnostic |
kaggle/ |
kernel metadata and notebooks per milestone |
results/ |
measured JSON artefacts (M0 ceiling, profiles, M1 gate, M2 sweep, M3 pre-flight) |
tests/ |
187 unit tests |
Code: MIT. Two frozen third-party models are downloaded at run time and not
redistributed here: DINOv2-small (facebook/dinov2-small, Apache-2.0) and DC-AE
(mit-han-lab/dc-ae-f64c128-in-1.0-diffusers; the SANA f32 variant was tried in M0
only). The DC-AE code (mit-han-lab/efficientvit) is Apache-2.0; the weight
repositories on the Hub carry no licence field of their own, so their status is
inherited from the code release rather than stated explicitly. Nothing in this
repository — no features, latents or checkpoints — derives from those weights;
all training data is procedurally generated. SD-VAE was evaluated for M0 only and
rejected on token count before its licence became relevant. Method references: RTFM (World Labs), REPA (Yu et al.), Muon
(Jordan et al.), muP (Yang et al., microsoft/mup), diffusion forcing (Chen et al.),
Target-Bench (arXiv 2511.17792).
Developed by one person over three weeks with AI coding agents (Claude) under a written brief whose central rule is the ledger: no GPU run without a ticket, no ticket without a profile, no result without its budget.
@software{nanowm_2026,
title = {nanoWM: which training tricks pay off for a 40M-parameter world model in 20 GPU-hours?},
author = {Say43},
year = {2026},
url = {https://github.com/Say43/World-Model}
}