Guidance for AI coding agents (and humans) working in this repository. See
ARCHITECTURE.md for the technical map of what each folder is and how they
relate; this file is conventions and how-to-run.
A staged model-training workspace — pre-training/ (data prep) →
fine-tuning/ (LoRA adapters) → serving/ (inference) → training/
(from-scratch training) — scaling from a single RTX 3090 (24GB) up
to multi-GPU. Each leaf folder is an independently deployable uv project;
there is no shared root Python environment. Full detail in
ARCHITECTURE.md.
uv only, one binary at the root drives every pipeline via
uv run --directory <folder> <command> — see the root
README.md for the current full list. Never cd
into a folder and run bare python; always go through uv run (or the
folder's uv_setup.bat / uv_bootstrap.bat for a one-time install) so the
correct pinned interpreter and CUDA torch build are used.
Each fine-tuning/serving folder pins its own torch/CUDA index in its
pyproject.toml — do not add a root-level pyproject.toml with shared
dependencies, and do not try to unify these into one uv workspace/lockfile.
The pins differ on purpose (different model classes, different CUDA
requirements) and a shared resolution would fight that.
Each project also pins its own .python-version (currently 3.12
everywhere). Don't remove these or let a project fall back to "whatever
Python uv finds newest" — a too-new CPython (e.g. 3.14) can lack prebuilt
wheels for pinned deps like pillow, which makes uv run fail trying to
build from source instead of installing a wheel.
fine-tuning/ currently holds vicuna-7b-lora/ and qwen25-3b-lora/, both
same transformers+peft pattern. An earlier Axolotl-based
axolotl-ocr-summary/ pipeline was removed by the repo owner — if a similar
pipeline reappears, note that axolotl[deepspeed] only resolves its uv
environment on Linux/WSL (triton ships no Windows wheels), a real platform
constraint, not something to patch around silently.
-
One root
.gitignorecovers the whole repo — don't add per-folder.gitignorefiles. Its patterns are unanchored on purpose so they match at any depth (DATASET/,data/,runs/,output/,outputs/,.cache/,hf_cache/,merged_model/,*.safetensors,*.pt,*.csvetc. are ignored everywhere, not per-project). Only code, configs,.python-version,uv.lock, and README stubs in those drop-zone folders are versioned. Datasets are downloaded/pointed-to locally, never checked in. Likewise: one rootAGENTS.mdfor the whole repo, not one per pipeline folder. -
fine-tuning/vicuna-7b-lora/builds its JSONL from CNN/DailyMail only (build_vicuna7b_dataset.py --cnn-dailymail-dir, required) — the earlier dual-source mode that also read pre-training's image-linked OCR/SUMMARIES CSV pair was removed entirely (not renamed) since it isn't needed for this pipeline's current use. Its CLI flags and JSONL field are named generically (--source-csvon the generator's batch-eval mode,--text,--text-file, JSONL fieldtext), notocr_*— don't reintroduce OCR-specific naming or resurrect the removed CSV-pair ingestion path without being asked. Its default instruction wrapper matches the CNN/DailyMail prompt, so--instructiondoesn't need to be passed explicitly for the common case; it's still a CLI flag (on both the trainer and generator, must match between the two) for training on differently-worded source text. -
Judge a trained
vicuna-7b-loraadapter bygenerate_vicuna7b_lora.py --jsonl-eval data/vicuna7b_train.jsonl --num-samples N, not loss alone. It replicates the trainer's held-out split and prints source/reference/ generated triples with token-F1 — loss can plateau while the model is still producing good summaries (verified on a predecessor run: a flat ~1.0–1.2 loss plateau still produced coherent, on-topic, correctly-styled summaries). -
vicuna-7b-loraloadslmsys/vicuna-7b-v1.5directly viaAutoModelForCausalLM/AutoTokenizer— not a LLaVA checkpoint, notLlavaForConditionalGeneration/AutoProcessor. Don't reintroduce a LLaVA dependency here; a plannedllava15-full-lorasibling (image+text pairs, actually exercising the vision encoder/projector) is where that belongs — this repo's first real VLM fine-tune.protobufis a required dependency in this pipeline specifically becauselmsys/vicuna-7b-v1.5ships a raw SentencePiece tokenizer that needs it to convert to a fast tokenizer. -
fine-tuning/qwen25-3b-lora/isvicuna-7b-lora's sibling, same pattern,Qwen/Qwen2.5-3B-Instructinstead. LoRAtarget_modulesstay["q_proj", "v_proj"](verified same as Vicuna viapeft's default LoRA target-module table forqwen2) but the prompt wrapper is ChatML (<|im_start|>role\n...<|im_end|>), not Vicuna'sUSER:/ASSISTANT:— verified against the model's actualtokenizer_config.json(eos_token="<|im_end|>") before writing the code, not assumed. Before cloning this pattern to a new base model, check its actual attention module names first —microsoft/Phi-3.5-mini-instruct, for example, fuses Q/K/V into a singleqkv_projlinear layer (confirmed by readingPhi3Attention's source), sotarget_modules=["q_proj","v_proj"]would silently attach to nothing on that model. -
training/imdb-sentiment-cnn/trains a Text CNN (Kim 2014) from scratch on the Large Movie Review Dataset — binary sentiment classification, 25k train / 25k test, judged by accuracy on the held-out 25k test split. Hand-written philosophy like the rest oftraining/: no torchtext/transformers/nltk, noDataLoader(numpy-permutation batching), and no GloVe/pretrained embeddings by design — strictly IMDB-only data, random-init trainable embeddings; don't silently add pretrained vectors.build_imdb_dataset.pyverifies exactly 12,500 review files per split and refuses to build on a partial extraction (the dataset was once caught mid-extraction withtrain/posstill filling up); a nestedaclImdb/subfolder is also accepted. Verified real run: 89.2% test acc, 20 epochs in ~31 s on the RTX 3090; dropout 0.5, best checkpoint by val acc (peaks ~epoch 3, then the model overfits fast — train acc → 100%); a dropout-0.7 variant scored worse and was discarded. The 50k unlabeled reviews are deliberately unused — a future AWD-LSTM / transformer+MLM pipeline is where they belong. -
training/flow-matching-mnist/trains a flow-matching / rectified-flow generative model from scratch on MNIST — the contemporary counterpart totraining/mnist-vae(same dataset, samedata/mnist.npzcontract, same hand-written zlib PNG writer, so the two sample grids are directly comparable). Hand-written philosophy like the rest oftraining/: nodiffusers/torchcfm/torchdiffeq/torchvision, noDataLoader(numpy-permutation batching) — the UNet velocity field, sinusoidal time embedding, EMA, and Euler/Heun ODE samplers are all written out. The objective ismse(v(x_t, t), x1 - (1-sigma_min)*x0)on the conditional-OT pathx_t = (1-(1-sigma_min)*t)*x0 + t*x1(Lipman et al. 2210.02747); at the default--sigma-min 0.0that is exactly rectified flow (Liu et al. 2209.03003). There is no noise schedule and no ELBO here on purpose — don't add betas/alpha_bar, a variance head, or loss reweighting; that turns it back into a DDPM. 1,175,841 params at--base-channels 32; same MNIST drop asmnist-kmeans/mnist-vae. The evaluator's round-trip MAE/PSNR sweep measures ODE discretization error, not sample quality — a near-zero velocity field round-trips almost perfectly since the identity is its own inverse, and a 2-epoch smoke run really did beat the converged model on it. Judge samples bysamples_grid.pngplus the nearest-neighbour memorization check. There is deliberately no FID (it needs a pretrained Inception, against this folder's from-scratch rule) — don't add a substitute score. -
training/rvq-audio-codec/trains a neural audio codec with residual vector quantization from scratch on LJSpeech — the EnCodec/SoundStream/DAC architecture, and the first audio pipeline in this repo. Hand-written philosophy like the rest oftraining/: the RIFF/WAVE parser and writer, the SEANet conv encoder/decoder, the RVQ (factorized 8-dim lookup, cosine distance, EMA updates, dead-code revival), the mel filterbank, the multi-scale STFT discriminator and SI-SDR are all written out — noencodec/descript-audio-codec/audiocraft, notorchaudio/librosa/soundfile/scipy, noDataLoader(numpy-permutation batching over utterance indices, one random crop each). 7,338,658 params plus a 2,112,582-param discriminator used only in training.--data-dirpoints at an extracted LJSpeech-1.1 (a nestedLJSpeech-1.1/subfolder is also accepted);build_ljspeech_dataset.pyverifies exactly 13,100 wavs and refuses a partial extraction, and writes a memmappeddata/ljspeech_audio.i16+data/ljspeech_index.npzrather than an.npzof samples — 3.8 GB of int16 becomes 7.6 GB as float32 and won't sit in RAM. Things not to silently undo:- Trained at LJSpeech's native 22,050 Hz, no resampler. Frame rate is 68.9 Hz and the bitrate 5.51 kbps — don't "fix" these to EnCodec's published 24 kHz / 75 Hz / 6 kbps figures, and don't add a resampler without being asked.
- Quantizer dropout is load-bearing, not a regularizer: it is what lets one trained model serve the whole 1→8 codebook ladder. Removing it means eight separate runs for the same demo.
- The discriminator is staged behind
--adv-start-stepon purpose;--lambda-adv 0is the reconstruction-only A/B. Don't make it unconditional. - The dead-code cutoff is a fraction of uniform codebook usage, not an
absolute count. The absolute 2.0 that EnCodec and
vector-quantize-pytorchuse is a trap at this batch size (32 × 69 = 2,208 vectors over 1,024 entries → uniform usage is 2.16), and the first smoke run really did report 1,023 of 1,024 entries revived per codebook. After the fix, 0. - The discriminator is ~8x the cost of the whole rest of the step
(7.75 steps/s reconstruction-only vs 0.95 fp32 with it, RTX 3090, batch
32).
--disc-bf16 1(default) runs only the critic under bf16 autocast → 2.01 steps/s, 10.8 GiB instead of 16.8, no architecture change. Don't extend that autocast over the generator: the codebook lookup and EMA updates must stay fp32, since bf16 EMA statistics stop accumulating small updates — the exact thing dead-code revival exists to detect.cudnn.benchmarkand TF32 were measured and do nothing here. - The EMA weights are only better once converged. After n steps they
still carry
ema_decay**nof the random init — the 802-step smoke checkpoint at decay 0.999 is 45% initialization and scored 9.02 mel against the live weights' 7.42. A full run leaves that behind (0.999**24060 = 4e-11), so--use-ema 1stays the default, butevaluate_codec.pycomputes the share from the checkpoint and warns above 1%. Don't remove that warning or theglobal_step/ema_decayfields it reads. - SI-SDR is a weak proxy for a GAN-trained codec — the adversarial
loss trades waveform/phase alignment for perceptual realism, so a
better-sounding model can score worse. Judge by the emitted
original_NN.wav/recon_nq{8,4,2,1}_NN.wavpairs and the per-codebook usage table. There is deliberately no ViSQOL/PESQ/ NISQA (external binary or pretrained network), the same rule that keeps FID out offlow-matching-mnist— don't add a substitute score. It is the deliberate successor oftraining/cifar10-vqvae(one codebook → eight; full-latent L2 lookup → 8-dim factorized cosine lookup), and its collapse-mitigation findings are meant to transfer back there — the 60-epoch run ended with all 1,024 entries of all eight codebooks in use, the deepest codebook carrying the highest perplexity of the stack (903.6 vs the first's 792.0). Don't re-run training to "improve" the numbers without being asked: the run cost 2 h 54 min and the repo owner has said it is finished. Its measured results are pinned in the pipeline README's "Verified runs" and inARCHITECTURE.md; treat them as the record.
-
training/fashion-mnist-dcgan/trains a DCGAN (Radford et al. 2015) from scratch on Fashion-MNIST - the repo's first GAN pipeline. Hand-written philosophy like the rest oftraining/: the generator and discriminator conv nets, theN(0, 0.02)weight init, one-sided label smoothing and the balanced D/G update loop are all written out - notorchvision, nokagglehub/pytorch-gan-metrics, noDataLoader(numpy-permutation batching).--data-dirpoints at a folder of IDX ubyte files (the Kaggle CSVs are also accepted);build_fashion_mnist_dataset.pyverifies exactly 60,000/10,000 and refuses a partial extraction, writingdata/fashion_mnist.npz(float32[0,1]plus class names; the trainer rescales to[-1,1]for Tanh output). Things not to silently undo:- 28x28 does not divide cleanly down DCGAN's canonical 32x32 ladder -
three stride-2 convs take 28 -> 14 -> 7 -> 3, so the discriminator's
last feature map is 3x3 (final 3x3 conv to a logit) and the generator
starts from a 7x7 grid, not 4x4. The shapes in
train_dcgan.pyare the verified ones; "fixing" them to the paper's 32x32 figures breaks the tensors. - A GAN is judged by its samples, not its loss. D/G losses move
adversarially and say little about quality, so
train_dcgan.pysaves a fixed-z sample grid every--sample-everyepochs (collapse becomes visible across training) andevaluate_dcgan.pyemitssamples_grid.png(hand-written zlib PNG writer, same asmnist-vae/flow-matching), the nearest-neighbour memorization check (L2 to the closest training image vs a real-image control) and a pairwise-diversity probe. There is deliberately no FID/IS - both need a pretrained Inception, the same rule that keeps FID out offlow-matching-mnistand ViSQOL out ofrvq-audio-codec. New pipeline - no verified-run numbers yet; pin them in the pipeline README's "Verified runs" once it has been run on the RTX 3090.
- 28x28 does not divide cleanly down DCGAN's canonical 32x32 ladder -
three stride-2 convs take 28 -> 14 -> 7 -> 3, so the discriminator's
last feature map is 3x3 (final 3x3 conv to a logit) and the generator
starts from a 7x7 grid, not 4x4. The shapes in
-
training/vit-cifar10/trains a Vision Transformer (Dosovitskiy et al. 2021, pre-LN / norm-first layout as popularized by DeiT) from scratch on CIFAR-10 - the repo's first attention-based vision model and its first from-scratch transformer of any kind. Hand-written philosophy like the rest oftraining/: the patch embedding, learned CLS token + positional embeddings, the pre-LN transformer blocks, and the multi-head self-attention (QKV projections, scaled dot-product, output projection) are all plaintorch.nn- notransformers/timm/torchvision, noDataLoader(numpy-permutation batching).build_cifar10_dataset.pyis the same stdlib-pickle parser / samedata/cifar10.npzcontract astraining/cifar10-vqvae. Things not to silently undo:- Flip+crop augmentation is plain torch ops, applied per batch in the
training loop (
torch.flip, 4px zero-pad + random crop, per-channel normalize with the hardcoded CIFAR-10 train statistics).--no-augmentis the documented A/B (measured 66.8% top-1 with aug at 60 epochs; the no-aug leg is still unmeasured), not a debug flag to remove. - The LR schedule is linear-warmup-then-cosine, not plain cosine.
ViTs train unstably from scratch without the warmup; the rest of
training/'s plainCosineAnnealingLRis deliberately not reused here. Don't "simplify" it back. - AdamW with weight decay 0.05, not plain Adam - the ViT default,
unlike the other torch trainers in this folder. Same reason.
Defaults are ~10.7M params (
--dim 384 --depth 6 --heads 6 --mlp-ratio 4), ~23 min for 60 epochs fp32 on the RTX 3090; best checkpoint by val acc on a 10% holdout, the 10k test split stays unseen untilevaluate_vit.py. The evaluator reports test top-1/top-5, per-class accuracy + confusion matrix, and writespredictions_grid.png(a hand-written zlib RGB PNG - green border = correct, red = wrong - no imaging library). There is deliberately no pretrained-feature score, the same rule that keeps FID out offlow-matching-mnistand ViSQOL out ofrvq-audio-codec. The patch-embed/block stack is the planned encoder for a future I-JEPA-style self-supervised pipeline. Verified real run on the RTX 3090 (repo owner): 66.82% test top-1 / 97.32% top-5 at the 60-epoch defaults in ~23 min (10,695,562 params, best val 67.50% at epoch 60, frog 81.4% / cat 44.7% per-class). That is below the ~80-86% figure this section's earlier draft expected - too optimistic for flip+crop-only at 60 epochs; the measured 66.8% is the record, and the correction is documented in the pipeline README.
- Flip+crop augmentation is plain torch ops, applied per batch in the
training loop (
-
training/mae-cifar100/trains a Masked Autoencoder (He et al. 2022) from scratch on CIFAR-100 - the repo's first representation-learning (self-supervised) pipeline, and the mask-reconstruct sibling of the I-JEPA-style rung planned inARCHITECTURE.md. Hand-written philosophy like the rest oftraining/: the patch embedding, positional embeddings, pre-LN transformer blocks, hand-written multi-head self-attention, the random masking, and the lightweight decoder are all plaintorch.nn- notransformers/timm/torchvision, noDataLoader(numpy-permutation batching). The encoder is vit-cifar10's patch-embed/block stack copied in by hand (pipelines never import each other's code), defaulting to patch 2 -> 256 patches (64 visible at 75% masking;--patch-size 4gives the sibling's literal 64-patch config).build_cifar100_dataset.pyis the same stdlib-pickle style asvit-cifar10's builder, but CIFAR-100 ships one 50ktrainfile + one 10ktestfile (plusmeta), verified by exact count. Things not to silently undo:- Masking is per-sample fixed-count (a random permutation keeping the
first 75%-complement), not a per-patch Bernoulli - the paper's
scheme, so every image is masked at exactly
--mask-ratio. - The loss is MSE on the masked patches only, with per-patch-normalized
targets (subtract patch mean / divide patch std - the MAE trick that
stops the decoder collapsing to patch means).
--no-patch-normis the documented A/B, not a flag to remove. - Pretraining does not normalize pixels (flip+crop only, raw [0,1]): the reconstruction targets ARE the pixels. Normalization belongs to the linear probe, with the hardcoded CIFAR-100 train statistics.
- The decoder is pretraining-only (~1M of the ~11.8M params);
linear_probe.pydiscards it and reads frozen encoder features (mean-pooled patch tokens - MAE has no CLS token). - The judge is a hand-written linear probe (a linear head trained
from scratch on the model's own frozen features, SGD momentum + cosine
per the paper) - not FID/Inception, the same no-pretrained-features
rule as everywhere else in
training/. Probe features are precomputed offline (normalize-only inputs, no probe-time flip+crop - documented). - Pos-embed / mask-token are trunc-normal 0.02 initialized (the
paper's init) - unlike vit-cifar10, whose pos embed is left at zero.
Defaults: 11,766,540 params (10,750,848 encoder / 1,015,692 decoder),
~35 s/epoch on the RTX 3090 (smoke-measured) -> ~36 min for 60
epochs; best checkpoint by a deterministic full-image val reconstruction
MSE, the 10k test split stays unseen until
linear_probe.py, which reports test top-1/top-5, coarse (20 superclass) top-1, per-class + 100x100 confusion matrix,test_metrics.txtand a hand-written zlibprobe_grid.png. Verified real run on the RTX 3090 (repo owner): 25.56% test top-1 / 53.41% top-5 (coarse 37.89%) at the 60-epoch defaults in ~36 min (2,166 s) - recorded on the final epoch-60 checkpoint, which probes better than the best-val epoch-34 one (24.65% / 53.15% / 36.80%): the val recon curve bottomed at epoch 34 while the masked-MSE kept improving, so the late features are the better representation (a finding - don't assume the best-val checkpoint holds the best features; the comparison is deterministic and was reproduced). oak_tree 69.0% / bowl 1.0% per-class. That is below the ~30-45% figure this section's earlier draft expected - too optimistic for 60 epochs / 45k images without probe-time augmentation; the measured 25.56% is the record, and the correction is documented in the pipeline README.
- Masking is per-sample fixed-count (a random permutation keeping the
first 75%-complement), not a per-patch Bernoulli - the paper's
scheme, so every image is masked at exactly
-
training/dit-cifar100/trains a class-conditional Diffusion Transformer (Peebles & Xie 2022 - the architecture Sora is built on) from scratch on CIFAR-100 - the repo's first class-conditional generative transformer, natural big sibling offlow-matching-mnist(same conditional-OT flow-matching objective, now conditioned on the 100 real fine classes, with classifier-free guidance). Hand-written philosophy like the rest oftraining/: the patch embedding, the frozen 2D sincos positional embedding, the adaLN-Zero transformer blocks, the hand-written multi-head self-attention, the final unpatchify layer, the class embedding + null token, the conditional-OT probability path, the velocity-regression loss, the EMA, and the Euler ODE sampler are all plaintorch.nn- nodiffusers/torchcfm/torchdiffeq/transformers/timm/torchvision, noDataLoader(numpy-permutation batching).build_cifar100_dataset.pyis the same stdlib-pickle builder / samedata/cifar100.npzcontract asmae-cifar100. Things not to silently undo:- The objective is conditional-OT flow matching, not the DiT paper's
DDPM -
mse(v(x_t, t, y), u)on the Lipman et al. path (--sigma-min 0= rectified flow); no noise schedule, no ELBO, no reweighting. Don't add betas/alpha_baror a variance head; that turns it back into a DDPM. - CFG is trained with class dropout to null token 100
(
--cfg-dropout 0.1, the DiT paper's value); at sample timev_cfg = v_uncond + cfg*(v_cond - v_uncond), and--cfg-scale 1.0is plain conditional sampling (one forward per step). The null token is the extra embedding slot, not a 101st class. - The architecture follows the paper exactly: conditioning dimension = model dim (not 4x), 256-dim sinusoidal time embedding, final layer shift/scale without a gate, frozen sincos pos embed (the paper ablated learned ones worse), xavier Linears + zeroed adaLN/final output init. Don't "fix" the embedding widths to another reimplementation's 4x convention - that silently changes the architecture.
- Judge by
samples_grid.png(100 class-conditional samples, one per fine class, row-major),cfg_sweep.png(same latents at CFG scales 1.0-5.0), the nearest-neighbour memorization check vs a real-image control, and the test velocity MSE - theflow-matching-mnistrule. Deliberately no FID (pretrained Inception), the same rule that keeps FID out offlow-matching-mnistand ViSQOL out ofrvq-audio-codec. Defaults are 9,828,876 params (--dim 256 --depth 8 --heads 8 --patch-size 2, 256 tokens - the same token count as DiT-S/4 at 256 px), ~75 s/epoch fp32 on the RTX 3090 at batch 256 -> ~75 min for 60 epochs; best checkpoint by a deterministic-seed val velocity MSE, the 10k test split stays unseen untilevaluate_dit.py. Verified real run on the RTX 3090 (repo owner): 60 epochs in 4,521 s (~75 min), train velocity MSE 0.5174 -> 0.1695, best val 0.1842 at epoch 55, test velocity MSE 0.1887 on the held-out 10k (EMA weights); generated samples sit ~47% farther from the training set (mean L2 12.105) than real unseen test images (8.210) - no memorization, the same direction asflow-matching-mnist. The ~0.17-0.19 loss floor is the irreducible conditional variance of the velocity target, not a defect.
- The objective is conditional-OT flow matching, not the DiT paper's
DDPM -
-
training/librispeech-speaker-id/trains speaker embeddings from scratch on LibriSpeech - an ECAPA-TDNN (Desplanques et al. 2020) by default with the x-vector / TDNN it superseded (Snyder et al. 2018) as--arch xvector. The repo's first discriminative audio pipeline (every other audio pipeline reconstructs a waveform; this maps audio to a label) and its first metric-learning objective (everything else is cross-entropy or MSE; this learns a space where cosine distance is the quantity of interest). Hand-written philosophy like the rest oftraining/: the dilated conv blocks, squeeze-excitation, the Res2Net channel split, attentive statistics pooling, the AAM-softmax head, the speed-perturbation resampler, the EER/minDCF metrics and both plot renderers are plaintorch.nn/numpy - nospeechbrain/kaldi/sidekit/pyannote, notorchaudio/librosa/soundfile/scipy, noDataLoader. The log-mel filterbank is copied in by hand fromtraining/rvq-audio-codecand works at 16 kHz unchanged because it takessample_rateas an argument - don't copy that pipeline's 22,050-based band edges across with it. ~2,939,616 params at--arch ecapadefaults, 2,589,140 at--arch xvector(same--channels 256width;--channels 512is 4,454,868 and no longer a like-for-like A/B). Measured 30-epoch runs: 7.0 min ECAPA / 5.3 min x-vector, and the A/B inverted the expected answer - x-vector won the val EER that selects the checkpoint (0.0095 vs 0.0162) but lost the held-out-speaker EER that matters (0.0989 vs 0.0725). It fits the training speakers tighter and transfers worse. Don't report the in-training number as the result, and don't collapse the split. Things not to silently undo:- The index split is three-way, not a boolean: 0 = train, 1 = val
(held-out utterances of seen speakers, closed-set), 2 = unseen
(held-out speakers, open-set). LibriSpeech is partitioned by speaker
and
train.clean.100's 251 speakers have zero overlap withtest.clean/dev.clean(40 each, disjoint - verified), so a 251-way classifier cannot be evaluated on test-clean at all. Collapsing the split to a boolean silently makes one of the two numbers meaningless. - The headline metric is the open-set EER on split 2, not the
closed-set accuracy on split 1 - that split also selected the
checkpoint, so it is mildly optimistic and
eval_metrics.txtsays so. - AAM-softmax is what makes the embedding a metric; a plain softmax head can score well on 251-way identification while producing useless embeddings. Don't swap it for cross-entropy to "simplify".
- The head runs in fp32 outside the bf16 autocast - its sqrt/where/one_hot path is not autocast-safe and it is ~5k params.
- Evaluation crops are centred and deterministic (
rng=Noneinsample_batch), not random; a random eval crop makes the val EER jitter more than the training signal. - Speed perturbation is augmentation, not dataset resampling - it never touches the val/unseen paths, and LibriSpeech is used at its native 16 kHz.
ffmpegis a hard dependency of the builder (FLAC's Rice-coded, bit-serial residuals make a pure-Python decoder over 57 GB impossible, not merely slow; precedent ispre-training/exec_1.bat->poppler), andpyarrowreads the parquet container (same category as the stdlibpicklethe CIFAR builders parse). The decode is batched through ffmpeg's concat demuxer and sliced by FLAC STREAMINFO sample counts, asserted against ffmpeg's byte count - don't drop that check.- The builder verifies each split against its canonical LibriSpeech
size (28,539 for train.clean.100) like
build_ljspeech_dataset.py's 13,100-wav check. Don't relax it. - No pretrained speaker-verification score - same rule that keeps FID
out of
flow-matching-mnistand ViSQOL out ofrvq-audio-codec. It is the warm-up rung for a planned CTC-ASR pipeline (same builder, same memmap contract, same front-end), and its natural join withrvq-audio-codecis speaker-conditioned codec training.
- The index split is three-way, not a boolean: 0 = train, 1 = val
(held-out utterances of seen speakers, closed-set), 2 = unseen
(held-out speakers, open-set). LibriSpeech is partitioned by speaker
and
-
training/mamba2-tinystories/is the repo's first non-attention sequence model: a Mamba-2 selective state-space model (the SSD / state-space-duality form) trained from scratch on TinyStories, with a parameter- and token-matched causal transformer as--arch transformerfor the A/B. Every earlier sequence model here is attention (ViT/MAE/DiT) or convolution (TextCNN/ECAPA); this is the first recurrence. Hand-written like the rest oftraining/: the input-dependent discretisation (dA = exp(dt*A),dB = dt*B), the chunked SSD scan, the causal depthwise conv, the SiLU-gated output projection, the tied-embedding LM head and a byte-level tokenizer are all plaintorch.nn- nomamba-ssm, nocausal-conv1d, notransformers/tokenizers, noDataLoader. Things not to silently undo:- The block gates with
out_proj(y * silu(z))and applies NO RMSNorm to the gate. Mamba-2's paper block puts anRMSNorm(gate)beforeout_proj; this one does not. That is a documented departure from the paper, not an oversight - an earlier draft of this entry wrongly described the block as having a "gated RMSNorm" when neither the code nor the task brief had one. Adding it is an architecture change, so the A/B would have to be re-run rather than the two being compared. - The tokenizer is the byte-level identity (vocab 256) on purpose. One token is one byte, so bits-per-token == bits-per-byte and the A/B is a clean comparison. Don't swap in BPE or a pretrained tokenizer without redoing the A/B; the point is that both legs see identical tokens.
- The chunked scan is asserted against a naive per-timestep reference.
--selftest(run once at startup;--no-selftestto skip) comparesmamba2_ssdagainstmamba2_scan_reference, and the assertion is the point - the naive one is an oracle, not a second implementation to maintain. Measured: fp32 forward relative error ~3.8e-06, fp64 ~6.6e-15. Independently reproduced against a clean-room oracle (exact 0.0 agreement for the reference itself). mamba2_ssd's output buffer must be allocated atcompute_dtype, not hard-coded float32. It was float32 first, which silently made the fp64 self-test column a second fp32 run - the fp64 forward error sat at 2.5e-08 (float32-level) instead of ~6.6e-15. A "control" that is not a control is worse than no control.- The canonical TinyStories counts are SEPARATOR counts, not story
counts. 2,119,718 and 21,989 are the number of
<|endoftext|>markers; stories = spans minus empty spans (train: 2,119,719 spans - 230 empty = 2,119,489; valid: 21,990 - 0 = 21,990; index total 2,141,479). The builder assertedstories = separators + 1first and REFUSED TO BUILD, which is the guardrail working. Don't widen the tolerance to make a wrong count pass, and don't "simplify" the accounting - the index carriestrain_separators,train_empty_spans,train_stories, ... precisely so this is auditable without rescanning 2 GB. - The corpus is a memmap (
data/tinystories_train.u16plus an index.npz), thervq-audio-codec/librispeech-speaker-idcontract, not an.npzof tokens: 1.94 G tokens is 3.9 GB as uint16. - The documented 10-minute config has been run to completion on an idle
RTX 3090 (§3b of the pipeline README), both legs, 1,104 steps x 32,768
tokens = 36.2 M tokens (1.91% of one pass): mamba2 3,515,008 params, 379 s
(0.343 s/step), best val 0.8004 (ppl 2.23); transformer 3,527,488 params
(+0.355%), 235 s (0.213 s/step), val 1.0053. The evaluator re-scores
them at 0.8197 vs 1.0002, the transformer's linearly-interpolated
absolute positional embeddings collapse past the trained L=512 (1.63 / 2.49
nats at 1024 / 2048 vs mamba2's 0.91 / 0.85), and the MQAR probe is at
chance for both (~0.002 vs chance 0.0020). It is still not a finding:
1.91% of a pass cannot separate two architecture families, and the
librispeech-speaker-idreversal is the precedent. The--dim 51230-minute config remains an extrapolation. - Every s/step figure for the three pipelines added in this pass was
measured with the other two resident on the same RTX 3090 (they were
trained concurrently), one of which was also crashed mid-probe by
CUBLAS_STATUS_INTERNAL_ERRORunder that contention. Treat all of those timings - includingtrain_nerf.py's ~30 ms/step - as pessimistic upper bounds on step time rather than clean-GPU numbers, and re-measure on an idle GPU before quoting an extrapolated wall clock as a specification.
- The block gates with
-
training/meanflow-cifar10/is the repo's first 1-NFE generative model: MeanFlow (average-velocity flow matching; Geng, Deng, Bai, Kolter, He 2025) trained from scratch on CIFAR-10 - the direct successor offlow-matching-mnistanddit-cifar100, same conditional-OT path, now regressing the average velocity so ONE network evaluation generates a sample. Hand-written: the class-conditional DiT block stack, the two time embeddings (tandr), the adaLN conditioning, the stop-gradient target, the EMA and the 1..N-step sampler; the JVP istorch.func.jvp, i.e. torch's own forward-mode autodiff, not a library method. Nodiffusers/torchcfm/torchdiffeq. Things not to silently undo:- The
vin the training target is the CONDITIONAL velocityv_cond = x1 - x0, NOT the model's own output. This is the one that really matters. The first version usedu_theta(z_t, t, t)forv, and that makesu_theta == 0an EXACT global optimum - if u=0 then v=0, the JVP=0, the target=0 and the loss=0 - so it "converges" to zero loss while learning nothing. Measured on a 2-D toy: with the model's own output the loss diverges at lr 2e-3, and at lr 1e-4 converges to a model whose sample MMD equals raw noise's; withv_condit trains and 1-NFE sampling genuinely works. The paper's Eq. 9-11 is explicit that the conditional velocity is the target's only ground-truth signal. The same bug was left behind inevaluate_meanflow.py, which built its "training objective" metric the same wrong way and therefore reportedtest average-velocity MSE 0.0000- a self-referential zero that ANY model scores, including a random one, and which reads as "solved". Corrected tov = x1 - x0, the same smoke checkpoint reports 1.2536, which is||v_cond||^2for CIFAR-10 in [-1,1]: with u ~ 0 the Jacobian ~ 0, the JVP term vanishes and the average-velocity MSE must equal the instantaneous one. Both metrics agreeing exactly is the predicted signature. Treat a near-zero average-velocity MSE on an undertrained model as a bug signal, not a result. - Two diagnostics are load-bearing evidence, not decoration. At
initialization the raw MSE must be O(1) ~
||v_cond||^2- measured 6.33 withv_condagainst 6.3e-04 with the model's own output - so a near-zero loss in the first steps means the degenerate target is back. And forcingr == tmust reduce EXACTLY to plain Flow Matching (measured identical to 1e-9, both 3.19850540), which the paper states and the trainer logs. - Stop-gradient on the target is essential (
.detach()): it is what avoids double backpropagation through the JVP. - No noise schedule, no betas/alpha_bar, no ELBO, no variance head - that
turns it back into a DDPM, the same rule
flow-matching-mnistanddit-cifar100carry. - Judge by the sample grid and the NFE sweep, never by loss, and there is deliberately no FID (pretrained Inception). The NFE sweep must render the SAME latents at every step count, or the 1-step claim is not falsifiable. Print the pixel standard deviation beside the NFE agreement: a barely-trained network emits a near-constant field, and a near-constant field makes one big step and fifty small ones agree TRIVIALLY. On the smoke checkpoint pixel std is 0.315 at NFE=1 against ~1.0 for real CIFAR-10 in [-1,1], so the small agreement number there is an artifact of the control, not evidence.
- The 1-NFE claim is UNMEASURED on CIFAR-10. It holds only on the 2-D toy
(
--selftest): NFE=1 MMD^2 0.01388 vs 0.00534 at NFE=20, against a raw-noise reference of 0.06551. The first real run (fast--dim 128 --depth 4config, 20 epochs / 7,040 steps in 1,258 s, test average-velocity MSE 0.6705 / instantaneous 0.2210, NN check 9.613 vs 9.080 real) does not change that: its NFE=1 samples still carry a pixel std of only 0.1746 against ~1.0 for real CIFAR-10, so its 0.15897 NFE=1-vs-50 agreement is not evidence either. Settling it needs the--dim 256 --depth 660-epoch default, now measured at ~10.2 h (0.867 s/step at batch 64; batch 256 exceeds the 24 GB card), which has not been run. Do not upgrade the toy or fast-config result to a CIFAR-10 1-NFE claim. - Do NOT copy the siblings' 1000x time-embedding scale into this
pipeline.
flow-matching-mnistanddit-cifar100scale the sin/cos time argument by 1000; MeanFlow cannot, because(t-r) * d/dt usits INSIDE the objective and multiplies the very time-derivative being learned. Measured after one epoch at lr 1e-4, raw average-velocity MSE: scale 1000 -> 35.09 (RISING from 1.32, i.e. diverging), 100 -> 1.014, 10 -> 0.821, 1 -> 0.812. Default is--time-scale 100.0, recorded in the checkpoint (the evaluator rebuilds with it, falling back to 1000 for older checkpoints). Diagnosed by logging the terms:|(t-r)*dudt|grew 0.000 -> 2.017 while|v_cond|held at 1.164 - the network was inflating its own time-derivative. --learning-ratedefaults to 1e-4, notdit-cifar100's 1e-3; 3e-4 and 1e-3 blew up to ~70 and ~157 within 80 steps on this objective. And--adaptive-weightcannot detect that divergence - its loss is the bounded ratiomse/(mse+1e-3), which tends to 1 as the model diverges - so the divergence guard watches the RAWu_mse, not the loss.--t-samplingdefaults to logit-normal, a deliberate deviation from the task brief (measured better at NFE=1: 0.01388 vs 0.03739 uniform, against a 0.06551 noise reference).--selftestmust INHERIT that default rather than carry its own; a selftest with a stale default validated a sampling scheme training no longer ran, which was a real bug here.
- The
-
training/3dgs-nerf-synthetic/is the repo's first 3D / novel-view synthesis pipeline: 3D Gaussian Splatting (Kerbl et al., SIGGRAPH 2023) trained from scratch on the NeRF synthetic scenes, with a hand-written vanilla NeRF (train_nerf.py, Mildenhall et al. 2020) as the representation-vs-representation baseline. Hand-written like the rest oftraining/: the PNG decoder (stdlib zlib plus the five filter types), the camera conversion, the tile-based differentiable rasterizer, the SSIM, the densification/pruning and the plot renderers - nonerfstudio,gsplat,diff-gaussian-rasterization,tiny-cuda-nn,plyfile,Pillow,torchvision, noDataLoader. Things not to silently undo:- The shipped PNGs are 8-bit RGBA (colour type 6), not RGB. A decoder that only accepts colour type 2 fails on every image. Verified bit-exact against an independent decoder on 10 images across 5 scenes, exercising all five filter types.
- Frames are keyed by
(split, file_path), never by filename. train, val and test are three DIFFERENT camera orbits that reuse the same names (train/r_0,val/r_0,test/r_0); a filename-keyed join pairs a training image with a validation camera, trains anyway, and surfaces only as a bad reconstruction. Verified: frame i is bit-exactly frame i's image and carries the correct world-to-camera at the split boundaries. - The
test/folder holds 600 PNGs buttransforms_test.jsonlists 200. The other 400 arer_N_depth_0001.pngandr_N_normal_0001.png. Read the JSON; never glob the directory. - The camera chain is validated by two invariants, not by eye: the world
origin must project to the image centre, and the silhouette carve must keep
a coherent object volume - the latter being what pins down roll, the one
error no geometric invariant catches. Don't remove either; both still
"train" when the convention is wrong. Two different measurements of the
projection invariant exist and must not be conflated: the builder
asserts < 2.0 px and measured a max of 1.16e-02 px across all 8 scenes x
3 splits (most scenes <= 2.4e-03), while
train_nerf.py --self-checkmeasures 6.6e-05 px on a single scene at--downscale 4. Both are correct; they are different code paths at different resolutions. An earlier draft of this entry quoted the self-check figure as if it were the builder's 8-scene result. - The tile rasterizer is asserted against a brute-force composite whenever
the top-K candidate bound is inactive, forward AND gradients:
--self-checkmeasured forward 0.000e+00 (bit-identical) and gradient 1.492e-13 in fp64. This is the branch's correctness evidence and it caught a wrong tile-scatter reshape that produced a factor-scale 0.5 error. The top-K bound and the 3-sigma bounding-box cull are deliberate approximations and the unculled difference is REPORTED, not asserted on - don't invent a pass/fail threshold for an algorithm property. - A pure-PyTorch rasterizer cannot run at 800x800, so the default is
--downscale 2(400x400). The published 3DGS figures are at 800x800 and are not comparable; the resolution trade is documented, not hidden. - There is deliberately NO LPIPS (pretrained VGG/AlexNet), the same rule
that keeps FID out of
flow-matching-mnistand ViSQOL out ofrvq-audio-codec; ARCHITECTURE.md stated this for the 3DGS branch before it existed. - The opacity reset must LOWER opacities, and stay strictly above the prune
threshold.
reset_opacity()writesmin(opacity, --opacity-reset-value)(default 0.01);--min-opacity(0.005) is the prune line. The first version clamped opacities UP to--min-opacity— the same constant the prune pass tests — so a reset parked every weak Gaussian exactly on the prune line and the next densify pass deleted them en masse: measured 2,628 of 2,992 (87.8%) at the step-3,000 reset of a--downscale 2run, which never recovered its density (383-406 Gaussians for the remaining 9,000 steps, val PSNR 20.408 -> 15.51) and was abandoned.main()now refuses to start when--opacity-reset-value <= --min-opacity. Don't collapse those two constants back into one. - Verified run (2026-09-23, idle RTX 3090):
lego, 5,000 steps at--downscale 4— 48.6 ms/step, 242.9 s, 5,544 Gaussians, best val PSNR 19.748 dB; held-out test 19.449 dB / SSIM 0.7941 on all 200 test frames, train-vs-test gap -0.256 dB (no overfitting). The pre-fix run of the identical config (6,388 Gaussians) scored test 18.779 dB / 0.7802, so the opacity-reset fix is worth +0.67 dB / +0.014 SSIM; both runs are kept because the second is only meaningful against the first. Still not run at the paper's budget or resolution: 5,000 steps at 200x200 is not 30,000 at 800x800, there is no--downscale 1run, and onlylegohas been trained. train_nerf.pyreusesSceneData,psnrandssimfromtrain_3dgs.pyso both models are scored by identical code. It is coarse sampling only - the paper's two-stage hierarchical sampler is not implemented, which is its biggest quality gap. Raw NeRF PSNR is expected to sit well below 3DGS at a comparable budget; the gap is the result, not a defect. Note its step cost is set by--batch-rays x --n-samplesand NOT by--downscale(training samples a fixed ray count per step; resolution only changes evaluation cost). The 300-step smoke run (16.04 dB) is still its only number — a 20,000-step run was started and stopped unfinished, and the README's §5 says the larger budget remains unrun.
-
Cross-folder references use full relative paths from repo root, e.g.
fine-tuning/vicuna-7b-lora/README.mdreaches a sibling pipeline via../../<stage>/<pipeline>/. When moving, renaming or removing a pipeline folder, grep the whole repo for its old path (READMEs and code comments, not just imports — these pipelines don't import each other's code, but do reference each other's paths for the adapter/cache directories) before considering the move done. -
serving/currently holds no pipelines.serving/vicuna-7b-lora/(a FastAPI service over the Vicuna adapter) was removed by the repo owner, the same wayfine-tuning/axolotl-ocr-summary/was. If a serving folder returns, keep the boundary the old one had:serving/<pipeline>/never importsfine-tuning/<pipeline>/code — it only reads that pipeline's trained output directory (adapter or merged model), which is what makes serving independently deployable. -
Batch/PowerShell scripts follow the existing style: bootstrap the env first (
uv_setup.bat/uv_bootstrap.bat), resolve paths from the script's own location rather than assuming a cwd,exit /b 1on failure. -
OCR-style input formatting is load-bearing where it appears (e.g.
(newline)markers, garbled spellings inpre-training/'s output) — that's the real training distribution downstream pipelines were tuned against. When wiring in a clean dataset like CNN/DailyMail instead, don't silently reformat it to look like OCR text; keep it as its own data source and be explicit about which fine-tuning run used which source. -
Root
README.mdis user-owned — only edit it when explicitly asked to.ARCHITECTURE.mdand this file are the agent-maintained docs; keep them in sync with structural changes (new pipeline folder, moved folder, changed data contract) as part of the same change, not as a follow-up. -
pre-training/is currently out of scope for active work per repo owner — don't modify it unless asked.