feat(jetson): add Orin Nano bring-up foundation - #1
Draft
KarlTaylorKnight wants to merge 9 commits into
Draft
KarlTaylorKnight wants to merge 9 commits into
KarlTaylorKnight wants to merge 9 commits into
Conversation
Add a fail-closed capability probe for Jetson, CUDA, memory, and NVMe-backed model storage. Document the staged Orin Nano implementation plan and add focused tests for platform, storage, error, schema, and CLI behavior.
…ming (Task 2) examples/bench.py could not run on the torch backend: it called core.reset_peak_memory()/get_peak_memory(), which only MLX has, and its host timers never waited for queued CUDA work. This is the portable reporting increment of docs/plans/jetson-orin-nano.md (Task 2); no physical Orin run is claimed. * examples/benchmark_report.py: pure, stdlib-only report builder (schema_version 1): validated per-run counts and timings, integer-byte memory metrics with method/scope, null-with-reason convention, strict JSON (allow_nan=False), probe linkage by content SHA-256, atomic no-clobber write. * examples/benchmark_measure.py: benchmark-owned measurement adapter (MLX peak memory kept as before; torch CUDA allocator peaks with torch.cuda.synchronize at phase boundaries on an actual cuda device; explicit unavailable values on cpu/mps), process RSS peak per OS, a sampler thread for rss/anon/file/swap, git identity, nvpmodel power mode, checkpoint manifest and adapter identity without hashing weights. * examples/bench.py: --json-output PATH (new file only), --probe-json PATH, --rss-sample-interval; synchronize after reset, after prefill, after warmup and after the final timed step; same two-run greedy-warmup / sampled fixed-length protocol and human output (peak_active=n/a when unavailable); pre-load validation; framework imported lazily; engine and sampler closed in finally; failures exit non-zero and leave no report. * tests: test_benchmark_report.py, test_benchmark_measure.py, test_bench_cli.py (fakes only) and test_bench_backend.py (real backend namespace, no weights). pyproject: pytest pythonpath = ["."] so plain `pytest` collects the repo-root scripts/examples imports. * docs: benchmark protocol and schema in docs/nvidia.md; the revised plan replaces docs/plans/jetson-orin-nano.md with its C01-C16 justifications alongside; README warmup wording corrected (it was always greedy). Host checks only (Windows torch-CPU venv, WSL Linux, Python 3.9/3.11): no CUDA execution, no MLX, no Orin evidence. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Gate A: validated torch 2.14.0+cu130 (official aarch64 wheel, the GB10 combination) on L4T R39.2 / CUDA 13.2 / Python 3.12; edge0 installed --no-deps with no MLX on the device. The wheel warns sm_87 is outside its SASS list; execution evidence gates instead of the warning. New acceptance deliverables: * tests/conftest.py: --require-cuda option (fail, never skip, when CUDA is required but absent — verified with CUDA_VISIBLE_DEVICES=""). * tests/test_torch_cuda_smoke.py: torch-only CUDA smoke against hand-computed references (int4 gather_qmm, uint32 word view, RMSNorm, core ops). 6/6 pass on the device. * scripts/orin_reference_check.py: 32-token greedy CUDA generation vs torch-CPU teacher-forced logits, tolerances registered before the run. Result: 32/32 token choices identical, zero near-ties; the vocab-wide max-|Δlogit| bound of 1.0 was exceeded (2.06, deep-tail bfloat16 noise; ≤ 0.22 at every step's top-8 tokens) and is recorded as that bound's FAIL, not loosened retroactively. Short baseline (3 independent launches x 2 runs, --ntok 32 --warmup 0, weight cache/prewarm off, tegrastats alongside): decode 0.53-0.58 tok/s, prefill 11.3-15.9 s @ 37 tok, CUDA peak <= 1.00 GiB, peak RSS 5.3-5.8 GiB, no OOM. Raw evidence stays in the untracked run dir; sanitized summary in docs/nvidia.md, status in the plan. Also: guard mlx imports in test_sampling/test_streaming_math with importorskip (plain pytest now collects on hosts without MLX), and gitignore models/. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
Pure calculation in streaming/budget.py: observed available RAM (tightened by any process limit) minus itemized deductions — OS-growth reserve, allocator allowance, KV at the DECLARED context, whole-layer prefill transient, in-flight builds, pinned staging (0 until Task 5) — leaves a byte allowance that bounds the shared LRU and prefetch buffer. Payloads are priced from the checkpoint's safetensors header (bundle_bytes_from_entries), never from slot counts. Impossible profiles raise BudgetError before inference with the arithmetic shown; a calculated starved cache is rejected, never constructed as SharedExpertCache(0)/PrefetchBuffer(0) (those disable eviction). Enforcement the unbudgeted path leaves open: PrefetchBuffer.max_cap (ceiling on set_cap — bounds prefetch_all()'s silent growth to num_experts + 32) and LayerOptions.max_inflight (bounds QUEUED speculative prefetch builds; demand loads are never dropped). Both default to the historical semantics when no budget is active. Integration (8B engine): EDGE0_MEMORY_BUDGET=auto|<bytes> resolves at cache-construction time; EDGE0_BUDGET_CONTEXT declares the context (default 1024), priced by the new ModelConfig.kv_bytes_per_token (measured 1.1 MB/token for this tier). EDGE0_TORCH_WEIGHT_CACHE=1 is vetoed unless its measured ~4.1 GB fits the post-cache headroom. The resolved policy is recorded on the engine and in the bench report's caches group. On-device (Orin Nano 8 GB): auto resolves ~4.2 GB usable, keeps the tested 64/48 profile (real bundle ~1.30 MiB/expert), benchmarks identically to the unoptimized baseline (0.52 tok/s, same peaks, RSS 5.55 GiB in range); a 4096-token declaration and the weight cache are both rejected pre-inference with itemized errors. 25 tests in tests/test_memory_budget.py; edge0/streaming/__init__ now exposes backend-bound names lazily so the pure modules import (and test) on hosts with no backend. Full suite: 254 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
…opt-in), measured on the Orin The measured profile reshaped this task. The tested 64-slot shared LRU is smaller than the 8B per-step working set (23 layers x K=8 = 184 bundles): on the Orin baseline it NEVER hits — every token rebuilds ~184 bundles from the mmap. Fixing that alone (512 slots, 65% hits) moved end-to-end decode by ~nothing: decode is compute-bound (GPU 99%, per-call dequantization). Per this task's own gate the pinned-buffer / CUDA-stream slot pipeline is deferred until Task 6 shrinks compute, with the evidence recorded in the plan. Implemented instead, all opt-in and scheduling-only: * EDGE0_CACHE_SLOTS=<int>: override the REQUESTED LRU size; the Task 4 budget may still lower it; zero/negative rejected. * LayerOptions.predict_prefetch / EDGE0_PREDICT_PREFETCH=1: route each step's prerouter predictions into the existing bounded prefetch() for NON-staged layers (prod_k8 keeps staged decode off; consumption stays on the exact _get_bundles path — wrong/late predictions are demand loads, never staged zero rows). Prefetch buffer sized to one predicted step (owners x K), per-layer in-flight bound defaults to K, and the budget prices the in-flight transient as per-layer bound x producing layers. Measured (3 launches x 2 runs, baseline workload, budget active): decode 0.568-0.623 tok/s mean 0.594 vs baseline mean 0.552 (+7.7%); 85% hit rate; 2038/2063 predicted builds consumed, 0 wasted, prefetch_wait 0.0s; CUDA peak 1.58 GiB and RSS 5.5-5.8 GiB inside the resolved budget; token output bit-identical to the Task 3 reference. Kept off by default per Gate D while compute dominates. 10 new tests (stager routing with fakes, env-profile plumbing); 264 passed total. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
Measured per-op decode shares on the Orin (synchronized timers, unoptimized profile): gather_qmm 37% (69 calls/step), dense QuantizedLinear 26% (235 calls/step), rest 37% — both int4 dequantize-then-matmul paths qualify for Task 6. docs/plans/rtx6000-task6-handoff.md: what kernel work can be developed on a workstation GPU (setup, the measured profile, three candidate designs in rising effort order — torch _weight_int4pack_mm repack, batched expert bmm, custom fused kernel — deliverables with guarded dispatch/fallback and pre-registered tolerances, cross-compile for compute_87) and what must stay on the Orin (design decision, runtime loading, reference check, benchmark protocol, Task 5 re-measurement). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
Three opt-in paths against the Orin's measured decode profile (gather_qmm 37%, dense QuantizedLinear 26%, both dequantize-then-matmul), developed per docs/plans/rtx6000-task6-handoff.md on an RTX PRO 6000 with torch 2.14.0+cu130 -- the Orin's torch version. All are OFF by default, guarded so every unsupported case falls through to the reference implementation unchanged, priced for the Task 4 budget, and recorded in the bench report. Which one the Orin adopts, and any tok/s, is the Orin's decision: this commit claims parity and bytes, not speed. * EDGE0_QMM_BATCHED=1 -- gather_qmm dequantizes the distinct experts of a call once and runs one batched matmul over a padded per-expert slab instead of a python loop of dequantize+matmul pairs. Affine 4-bit only; 2/8-bit and any call whose priced transient exceeds EDGE0_QMM_BATCHED_MAX_BYTES (256 MiB) take the reference loop. Parity: 1e-4 to a float64 hand reference, 1e-5 to the loop, and within one bfloat16 ulp at the real decode shapes. * EDGE0_TORCH_WEIGHT_CACHE_BYTES=<n> -- keep the dequantized weights of the dense linears that fit a byte cap (WeightCachePolicy, first-fit in checkpoint order), shrunk by the engine to the budget's headroom. * EDGE0_INT4PACK=1 -- dense linears through torch's built-in _weight_int4pack_mm after a one-time repack of the MLX layout (uint32 LSB-first codes -> uint8 high-nibble-first, zero = bias + 8*scale). Approximate by construction; bounds registered in the tests before any Orin run. int4pack.probe executes the kernel once per module on the device first, so a wheel without an image for the target arch falls through instead of raising -- sm_87 is answered by execution, not by a table. The fill transient, not just the resident bytes: Whole-weight dequantization of edge0-8b's lm_head (157184 x 1536) peaks at 3702 MiB to produce a 460 MiB cache. A policy that admitted on resident bytes alone would have OOM'd an 8 GB Orin at the first forward after the budget said it fits. The fill is now chunked (measured 557 MiB, dequantized weight bit-identical), priced by nn.fill_transient_bytes, counted per module at admission and deducted by budget.dense_cache_request. EDGE0_TORCH_WEIGHT_CACHE=1 shares the branch and inherits the fix; it is now priced from the loader's measured total (235 cacheable linears, 1.402 GB bf16) rather than a hard-coded figure. Reference check: --test-env / --reference-env run the two phases with different knobs on the same device, and the report now carries host identity (GPU, compute capability, torch build, /proc/device-tree/model) plus each phase's RESOLVED knobs, so a workstation report cannot be read as an Orin one and a same-device run whose phases resolve identically warns that it proves nothing. Schema edge0-reference-check/2. Also fixes a real portability bug: the verdict line printed a non-ASCII character and raised UnicodeEncodeError on a cp1252 console AFTER writing the report, turning a PASS into a crash. Same-device reference checks on real weights, 32 greedy tokens, all keeping every token choice with zero divergences: the capped weight cache is bit-identical end to end (max abs logit diff exactly 0); the batched gather measures 1.01 and the int4 kernel 2.06, both exceeding the 1.0 bound Task 3 registered. Recorded as that bound's failure, not loosened -- as in Task 3, and the Orin's own check decides. A finding that reshapes this task's dense half: traced on the real model, 400 of 470 dense QuantizedLinear calls per step arrive with float32 activations and only 70 with bfloat16. Both dense knobs need bf16, so the cache fills 35 of the 235 modules it admits (0.174 GB of 1.402 GB reserved) and the int4 kernel repacks exactly those same 35 -- which is why neither moves decode. The reservation stays conservative (it never under-reserves) and the report now states filled_bytes beside admitted_bytes. The dense 26% is gated by activation dtype, not by the kernel; the batched gather covers the routed 37% and is unaffected. Tests across tests/test_quant_paths.py, tests/test_weight_cache_policy.py, tests/test_int4pack.py and tests/test_task6_budget.py: hand-computed references, no MLX, CUDA cases skipping without a device and failing under --require-cuda. Suite: 326 passed on the CUDA venv, 300 on torch-CPU, plus Python 3.9 and the WSL Linux portability subsets (the 14 test_jetson_probe failures are the known Windows POSIX fixtures). Design C, a custom fused kernel, was scoped to a cross-compile toolchain record only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…es the default Ran the acceptance the handoff reserved for this device: suite 341 passed on-device (96 under --require-cuda), same-device 32-token reference checks per path, and the 3-launch x 2-run benchmark protocol with the budget active. int4pack.probe EXECUTES on sm_87 — the kernel the wheel's SASS list does not advertise runs here, decided by execution as designed. Adopted: EDGE0_QMM_BATCHED defaults ON for this backend (EDGE0_QMM_BATCHED=0 selects the reference loop). It is the only candidate that both passed the registered reference bound on the target (32/32 token choices, max |dlogit| 0.777 vs the 1.0 bound) and gave a measured benefit there: decode 0.623-0.699, mean 0.668 tok/s, +20.7% over the 0.553 baseline, unchanged CUDA peak. The budget prices it with no env var set (kernel_transient 268,435,456 in the report). Gate D's bar — measured benefit before defaulting — is met on two GPU generations (+59% on the workstation). Not adopted, kept opt-in: the capped weight cache (+0.8%) and the int4 kernel (+3.7%); the activation-dtype gate limits both to 35 of 235 modules, and both exceed the registered bound (1.58 and 2.17), recorded as that bound's failures rather than loosened. Gate D's Task 5 re-measurement: the transfer knobs were worth +7.7% before this and are worth +17.5% ON TOP of the batched gather now (0.668 -> 0.785 tok/s; expert load wall 72.5s -> 17.0s over the same six runs). EDGE0_MEMORY_BUDGET=auto EDGE0_CACHE_SLOTS=512 EDGE0_PREDICT_PREFETCH=1 is the recommended Orin profile and stays opt-in as board-specific tuning. End to end: 0.785 tok/s, +42% over the Task 3 baseline, inside the budget, no OOM. Two workstation claims corrected, both only visible on the target: * the capped weight cache is NOT bit-identical on sm_87. Isolated directly: the cached weights ARE bit-identical, but one GEMM vs the reference's per-4096-row GEMMs differs by 0.03125 on a 30.6 output scale — one bf16 ulp of cuBLAS accumulation order. The exactness was a property of the workstation's heuristics, not of the code. * the float32 activation gate is deliberate MLX promotion parity, not an accident: layer 3 is the first BailingMLA, whose RoPE path computes in float32 and concatenates q_nope.to(q_pe.dtype) with a float32 q_pe, so the residual stream is float32 from there on and only layers 0-2 (KDA) present bf16 to their dense linears. Reaching the rest of the dense 26% is a numerics-parity decision for review, not a kernel increment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
…ated Bring-up plan complete. Everything measured on the target with the adopted profile (shipped default + EDGE0_MEMORY_BUDGET=auto EDGE0_CACHE_SLOTS=512 EDGE0_PREDICT_PREFETCH=1), desktop resident. * Repeat measurement, 3 launches x 2 runs: decode 0.744-0.786 tok/s, mean 0.772, sd 0.016 (n=6). Descriptive statistics of a small sample. * Correctness repeat: same-device reference check of the final profile PASSES the registered bound (32/32 tokens, max |dlogit| 0.777 vs 1.0); all 30 sustained requests produced byte-identical tokens. * Sustained 25 min / 30 requests / 960 tokens: mean 0.817 tok/s, sd 0.028, quartiles 0.791/0.834/0.816/0.824 — no degradation. Junction 54.5-61.8 C, NO throttling. RSS moved 5.45-5.70 GiB rising AND falling (mmapped expert pages reclaimed and re-faulted, not a leak); CUDA peaks flat at 1.57 GiB. * Swap marked explicitly: the run DID swap, +1165 MB in the first fifth then flat (+24/0/+18/-2), no throughput loss — a one-time eviction of the desktop's idle pages, not the working set. Per Gate C no swap-free resident-memory claim is made for this board. * Envelope: 3292-token prompt validated (prefill 42.4-42.8 tok/s, CUDA peak 3.14 GiB, no OOM); generation validated to 256 tokens (0.807/0.820 tok/s, no decay). kv_bytes_per_token for this tier: 1,100,000 -> 600,000. The old value was the README's Apple/MLX figure; measured here across 128..2048-token prefills the marginal cost is 561,320 B/token (548 KiB), about half. Growth is sublinear below the top of the range because most layers of this hybrid are KDA with fixed-size state and only the MLA layers' KV grows with context, so the top-of-range slope is the right constant to extrapolate from. Over-reserving was safe but rejected contexts that fit: the declared context the budget admits goes ~2560 -> 4096, which the 3292-token run then validated in practice. The claim, qualified to what was tested: edge0-8b runs on a physical Jetson Orin Nano 8 GB at 0.77-0.82 tok/s decode, 42 tok/s prefill at 3.3k tokens, +42% over the Task 3 baseline, within a budget that rejects what does not fit before loading. Single request, greedy sampling, <= 4096 declared context, <= 256 validated generation, 25W. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Why this branch is stacked
This work is based on upstream PR Edge0-AI#19 (
agourakis82:cuda-backend), not Edge0main. PR Edge0-AI#19 supplies the Torch/CUDA correctness reference and model ports that this Jetson effort needs. The proposed Orin work deliberately starts with measurement rather than another CUDA rewrite.Verification
uv run --no-project --with pytest python -m pytest tests/test_jetson_probe.py -q— 24 passeduvx ruff check scripts/jetson_probe.py tests/test_jetson_probe.py— passedpython3 -m py_compile scripts/jetson_probe.py tests/test_jetson_probe.py— passeduvx pyright scripts/jetson_probe.py— 0 errors, 0 warningsgit diff --cached --check— passedFull-suite limitation
The full suite could not be installed on the current x86_64 macOS host because the pinned
mlx==0.30.4package publishes macOS wheels for arm64, not x86_64. No physical Orin inference or performance claim is included in this PR.Next acceptance gate
Run the probe and existing CUDA smoke tests on an 8 GB Orin Nano with the Edge0-8B checkpoint on NVMe, then save separate prefill/decode, RSS, CUDA allocation, and thermal/power evidence before optimizing the data path.
Related upstream work