Skip to content

feat(jetson): add Orin Nano bring-up foundation - #1

Draft
KarlTaylorKnight wants to merge 9 commits into
cuda-backend-basefrom
feat/jetson-orin-nano-foundation
Draft

KarlTaylorKnight wants to merge 9 commits into
cuda-backend-basefrom
feat/jetson-orin-nano-foundation

Conversation

@KarlTaylorKnight

Copy link
Copy Markdown
Owner

Summary

  • adds a read-only Jetson Orin capability probe with stable JSON output
  • detects Jetson/L4T identity, shared-memory facts, CUDA/PyTorch status, and NVMe-backed model storage
  • returns explicit readiness blockers instead of claiming unsupported hardware works
  • documents a staged Orin Nano engineering plan with correctness, memory, and performance gates

Why this branch is stacked

This work is based on upstream PR Edge0-AI#19 (agourakis82:cuda-backend), not Edge0 main. PR Edge0-AI#19 supplies the Torch/CUDA correctness reference and model ports that this Jetson effort needs. The proposed Orin work deliberately starts with measurement rather than another CUDA rewrite.

Verification

  • uv run --no-project --with pytest python -m pytest tests/test_jetson_probe.py -q — 24 passed
  • uvx ruff check scripts/jetson_probe.py tests/test_jetson_probe.py — passed
  • python3 -m py_compile scripts/jetson_probe.py tests/test_jetson_probe.py — passed
  • uvx pyright scripts/jetson_probe.py — 0 errors, 0 warnings
  • git diff --cached --check — passed
  • CLI smoke on the development Mac correctly returned matching stdout/file JSON plus a non-zero status with explicit blockers

Full-suite limitation

The full suite could not be installed on the current x86_64 macOS host because the pinned mlx==0.30.4 package publishes macOS wheels for arm64, not x86_64. No physical Orin inference or performance claim is included in this PR.

Next acceptance gate

Run the probe and existing CUDA smoke tests on an 8 GB Orin Nano with the Edge0-8B checkpoint on NVMe, then save separate prefill/decode, RSS, CUDA allocation, and thermal/power evidence before optimizing the data path.

Related upstream work

KarlTaylorKnight and others added 9 commits September 13, 2026 17:58
Add a fail-closed capability probe for Jetson, CUDA, memory, and NVMe-backed model storage. Document the staged Orin Nano implementation plan and add focused tests for platform, storage, error, schema, and CLI behavior.
…ming (Task 2)

examples/bench.py could not run on the torch backend: it called
core.reset_peak_memory()/get_peak_memory(), which only MLX has, and its
host timers never waited for queued CUDA work.  This is the portable
reporting increment of docs/plans/jetson-orin-nano.md (Task 2); no
physical Orin run is claimed.

* examples/benchmark_report.py: pure, stdlib-only report builder
  (schema_version 1): validated per-run counts and timings, integer-byte
  memory metrics with method/scope, null-with-reason convention, strict
  JSON (allow_nan=False), probe linkage by content SHA-256, atomic
  no-clobber write.
* examples/benchmark_measure.py: benchmark-owned measurement adapter
  (MLX peak memory kept as before; torch CUDA allocator peaks with
  torch.cuda.synchronize at phase boundaries on an actual cuda device;
  explicit unavailable values on cpu/mps), process RSS peak per OS, a
  sampler thread for rss/anon/file/swap, git identity, nvpmodel power
  mode, checkpoint manifest and adapter identity without hashing weights.
* examples/bench.py: --json-output PATH (new file only), --probe-json
  PATH, --rss-sample-interval; synchronize after reset, after prefill,
  after warmup and after the final timed step; same two-run
  greedy-warmup / sampled fixed-length protocol and human output
  (peak_active=n/a when unavailable); pre-load validation; framework
  imported lazily; engine and sampler closed in finally; failures exit
  non-zero and leave no report.
* tests: test_benchmark_report.py, test_benchmark_measure.py,
  test_bench_cli.py (fakes only) and test_bench_backend.py (real backend
  namespace, no weights).  pyproject: pytest pythonpath = ["."] so plain
  `pytest` collects the repo-root scripts/examples imports.
* docs: benchmark protocol and schema in docs/nvidia.md; the revised plan
  replaces docs/plans/jetson-orin-nano.md with its C01-C16 justifications
  alongside; README warmup wording corrected (it was always greedy).

Host checks only (Windows torch-CPU venv, WSL Linux, Python 3.9/3.11):
no CUDA execution, no MLX, no Orin evidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Gate A: validated torch 2.14.0+cu130 (official aarch64 wheel, the GB10
combination) on L4T R39.2 / CUDA 13.2 / Python 3.12; edge0 installed
--no-deps with no MLX on the device. The wheel warns sm_87 is outside
its SASS list; execution evidence gates instead of the warning.

New acceptance deliverables:
* tests/conftest.py: --require-cuda option (fail, never skip, when
  CUDA is required but absent — verified with CUDA_VISIBLE_DEVICES="").
* tests/test_torch_cuda_smoke.py: torch-only CUDA smoke against
  hand-computed references (int4 gather_qmm, uint32 word view,
  RMSNorm, core ops). 6/6 pass on the device.
* scripts/orin_reference_check.py: 32-token greedy CUDA generation vs
  torch-CPU teacher-forced logits, tolerances registered before the
  run. Result: 32/32 token choices identical, zero near-ties; the
  vocab-wide max-|Δlogit| bound of 1.0 was exceeded (2.06, deep-tail
  bfloat16 noise; ≤ 0.22 at every step's top-8 tokens) and is recorded
  as that bound's FAIL, not loosened retroactively.

Short baseline (3 independent launches x 2 runs, --ntok 32 --warmup 0,
weight cache/prewarm off, tegrastats alongside): decode 0.53-0.58
tok/s, prefill 11.3-15.9 s @ 37 tok, CUDA peak <= 1.00 GiB, peak RSS
5.3-5.8 GiB, no OOM. Raw evidence stays in the untracked run dir;
sanitized summary in docs/nvidia.md, status in the plan.

Also: guard mlx imports in test_sampling/test_streaming_math with
importorskip (plain pytest now collects on hosts without MLX), and
gitignore models/.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
Pure calculation in streaming/budget.py: observed available RAM
(tightened by any process limit) minus itemized deductions — OS-growth
reserve, allocator allowance, KV at the DECLARED context, whole-layer
prefill transient, in-flight builds, pinned staging (0 until Task 5) —
leaves a byte allowance that bounds the shared LRU and prefetch buffer.
Payloads are priced from the checkpoint's safetensors header
(bundle_bytes_from_entries), never from slot counts. Impossible
profiles raise BudgetError before inference with the arithmetic shown;
a calculated starved cache is rejected, never constructed as
SharedExpertCache(0)/PrefetchBuffer(0) (those disable eviction).

Enforcement the unbudgeted path leaves open: PrefetchBuffer.max_cap
(ceiling on set_cap — bounds prefetch_all()'s silent growth to
num_experts + 32) and LayerOptions.max_inflight (bounds QUEUED
speculative prefetch builds; demand loads are never dropped). Both
default to the historical semantics when no budget is active.

Integration (8B engine): EDGE0_MEMORY_BUDGET=auto|<bytes> resolves at
cache-construction time; EDGE0_BUDGET_CONTEXT declares the context
(default 1024), priced by the new ModelConfig.kv_bytes_per_token
(measured 1.1 MB/token for this tier). EDGE0_TORCH_WEIGHT_CACHE=1 is
vetoed unless its measured ~4.1 GB fits the post-cache headroom. The
resolved policy is recorded on the engine and in the bench report's
caches group.

On-device (Orin Nano 8 GB): auto resolves ~4.2 GB usable, keeps the
tested 64/48 profile (real bundle ~1.30 MiB/expert), benchmarks
identically to the unoptimized baseline (0.52 tok/s, same peaks, RSS
5.55 GiB in range); a 4096-token declaration and the weight cache are
both rejected pre-inference with itemized errors.

25 tests in tests/test_memory_budget.py; edge0/streaming/__init__ now
exposes backend-bound names lazily so the pure modules import (and
test) on hosts with no backend. Full suite: 254 passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
…opt-in), measured on the Orin

The measured profile reshaped this task. The tested 64-slot shared LRU
is smaller than the 8B per-step working set (23 layers x K=8 = 184
bundles): on the Orin baseline it NEVER hits — every token rebuilds
~184 bundles from the mmap. Fixing that alone (512 slots, 65% hits)
moved end-to-end decode by ~nothing: decode is compute-bound (GPU 99%,
per-call dequantization). Per this task's own gate the pinned-buffer /
CUDA-stream slot pipeline is deferred until Task 6 shrinks compute,
with the evidence recorded in the plan.

Implemented instead, all opt-in and scheduling-only:
* EDGE0_CACHE_SLOTS=<int>: override the REQUESTED LRU size; the Task 4
  budget may still lower it; zero/negative rejected.
* LayerOptions.predict_prefetch / EDGE0_PREDICT_PREFETCH=1: route each
  step's prerouter predictions into the existing bounded prefetch()
  for NON-staged layers (prod_k8 keeps staged decode off; consumption
  stays on the exact _get_bundles path — wrong/late predictions are
  demand loads, never staged zero rows). Prefetch buffer sized to one
  predicted step (owners x K), per-layer in-flight bound defaults to
  K, and the budget prices the in-flight transient as per-layer bound
  x producing layers.

Measured (3 launches x 2 runs, baseline workload, budget active):
decode 0.568-0.623 tok/s mean 0.594 vs baseline mean 0.552 (+7.7%);
85% hit rate; 2038/2063 predicted builds consumed, 0 wasted,
prefetch_wait 0.0s; CUDA peak 1.58 GiB and RSS 5.5-5.8 GiB inside the
resolved budget; token output bit-identical to the Task 3 reference.
Kept off by default per Gate D while compute dominates.

10 new tests (stager routing with fakes, env-profile plumbing);
264 passed total.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
Measured per-op decode shares on the Orin (synchronized timers,
unoptimized profile): gather_qmm 37% (69 calls/step), dense
QuantizedLinear 26% (235 calls/step), rest 37% — both int4
dequantize-then-matmul paths qualify for Task 6.

docs/plans/rtx6000-task6-handoff.md: what kernel work can be developed
on a workstation GPU (setup, the measured profile, three candidate
designs in rising effort order — torch _weight_int4pack_mm repack,
batched expert bmm, custom fused kernel — deliverables with guarded
dispatch/fallback and pre-registered tolerances, cross-compile for
compute_87) and what must stay on the Orin (design decision, runtime
loading, reference check, benchmark protocol, Task 5 re-measurement).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
Three opt-in paths against the Orin's measured decode profile
(gather_qmm 37%, dense QuantizedLinear 26%, both dequantize-then-matmul),
developed per docs/plans/rtx6000-task6-handoff.md on an RTX PRO 6000 with
torch 2.14.0+cu130 -- the Orin's torch version. All are OFF by default,
guarded so every unsupported case falls through to the reference
implementation unchanged, priced for the Task 4 budget, and recorded in
the bench report. Which one the Orin adopts, and any tok/s, is the
Orin's decision: this commit claims parity and bytes, not speed.

* EDGE0_QMM_BATCHED=1 -- gather_qmm dequantizes the distinct experts of
  a call once and runs one batched matmul over a padded per-expert slab
  instead of a python loop of dequantize+matmul pairs. Affine 4-bit
  only; 2/8-bit and any call whose priced transient exceeds
  EDGE0_QMM_BATCHED_MAX_BYTES (256 MiB) take the reference loop.
  Parity: 1e-4 to a float64 hand reference, 1e-5 to the loop, and within
  one bfloat16 ulp at the real decode shapes.
* EDGE0_TORCH_WEIGHT_CACHE_BYTES=<n> -- keep the dequantized weights of
  the dense linears that fit a byte cap (WeightCachePolicy, first-fit in
  checkpoint order), shrunk by the engine to the budget's headroom.
* EDGE0_INT4PACK=1 -- dense linears through torch's built-in
  _weight_int4pack_mm after a one-time repack of the MLX layout
  (uint32 LSB-first codes -> uint8 high-nibble-first, zero = bias +
  8*scale). Approximate by construction; bounds registered in the tests
  before any Orin run. int4pack.probe executes the kernel once per
  module on the device first, so a wheel without an image for the target
  arch falls through instead of raising -- sm_87 is answered by
  execution, not by a table.

The fill transient, not just the resident bytes:

  Whole-weight dequantization of edge0-8b's lm_head (157184 x 1536)
  peaks at 3702 MiB to produce a 460 MiB cache. A policy that admitted
  on resident bytes alone would have OOM'd an 8 GB Orin at the first
  forward after the budget said it fits. The fill is now chunked
  (measured 557 MiB, dequantized weight bit-identical), priced by
  nn.fill_transient_bytes, counted per module at admission and deducted
  by budget.dense_cache_request. EDGE0_TORCH_WEIGHT_CACHE=1 shares the
  branch and inherits the fix; it is now priced from the loader's
  measured total (235 cacheable linears, 1.402 GB bf16) rather than a
  hard-coded figure.

Reference check: --test-env / --reference-env run the two phases with
different knobs on the same device, and the report now carries host
identity (GPU, compute capability, torch build, /proc/device-tree/model)
plus each phase's RESOLVED knobs, so a workstation report cannot be read
as an Orin one and a same-device run whose phases resolve identically
warns that it proves nothing. Schema edge0-reference-check/2. Also fixes
a real portability bug: the verdict line printed a non-ASCII character
and raised UnicodeEncodeError on a cp1252 console AFTER writing the
report, turning a PASS into a crash.

Same-device reference checks on real weights, 32 greedy tokens, all
keeping every token choice with zero divergences: the capped weight
cache is bit-identical end to end (max abs logit diff exactly 0); the
batched gather measures 1.01 and the int4 kernel 2.06, both exceeding
the 1.0 bound Task 3 registered. Recorded as that bound's failure, not
loosened -- as in Task 3, and the Orin's own check decides.

A finding that reshapes this task's dense half: traced on the real
model, 400 of 470 dense QuantizedLinear calls per step arrive with
float32 activations and only 70 with bfloat16. Both dense knobs need
bf16, so the cache fills 35 of the 235 modules it admits (0.174 GB of
1.402 GB reserved) and the int4 kernel repacks exactly those same 35 --
which is why neither moves decode. The reservation stays conservative
(it never under-reserves) and the report now states filled_bytes beside
admitted_bytes. The dense 26% is gated by activation dtype, not by the
kernel; the batched gather covers the routed 37% and is unaffected.

Tests across tests/test_quant_paths.py,
tests/test_weight_cache_policy.py, tests/test_int4pack.py and
tests/test_task6_budget.py: hand-computed references, no MLX, CUDA cases
skipping without a device and failing under --require-cuda. Suite: 326
passed on the CUDA venv, 300 on torch-CPU, plus Python 3.9 and the WSL Linux
portability subsets (the 14 test_jetson_probe failures are the known
Windows POSIX fixtures). Design C, a custom fused kernel, was scoped to
a cross-compile toolchain record only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…es the default

Ran the acceptance the handoff reserved for this device: suite 341
passed on-device (96 under --require-cuda), same-device 32-token
reference checks per path, and the 3-launch x 2-run benchmark protocol
with the budget active. int4pack.probe EXECUTES on sm_87 — the kernel
the wheel's SASS list does not advertise runs here, decided by
execution as designed.

Adopted: EDGE0_QMM_BATCHED defaults ON for this backend
(EDGE0_QMM_BATCHED=0 selects the reference loop). It is the only
candidate that both passed the registered reference bound on the target
(32/32 token choices, max |dlogit| 0.777 vs the 1.0 bound) and gave a
measured benefit there: decode 0.623-0.699, mean 0.668 tok/s, +20.7%
over the 0.553 baseline, unchanged CUDA peak. The budget prices it with
no env var set (kernel_transient 268,435,456 in the report). Gate D's
bar — measured benefit before defaulting — is met on two GPU
generations (+59% on the workstation).

Not adopted, kept opt-in: the capped weight cache (+0.8%) and the int4
kernel (+3.7%); the activation-dtype gate limits both to 35 of 235
modules, and both exceed the registered bound (1.58 and 2.17), recorded
as that bound's failures rather than loosened.

Gate D's Task 5 re-measurement: the transfer knobs were worth +7.7%
before this and are worth +17.5% ON TOP of the batched gather now
(0.668 -> 0.785 tok/s; expert load wall 72.5s -> 17.0s over the same six
runs). EDGE0_MEMORY_BUDGET=auto EDGE0_CACHE_SLOTS=512
EDGE0_PREDICT_PREFETCH=1 is the recommended Orin profile and stays
opt-in as board-specific tuning. End to end: 0.785 tok/s, +42% over the
Task 3 baseline, inside the budget, no OOM.

Two workstation claims corrected, both only visible on the target:
* the capped weight cache is NOT bit-identical on sm_87. Isolated
  directly: the cached weights ARE bit-identical, but one GEMM vs the
  reference's per-4096-row GEMMs differs by 0.03125 on a 30.6 output
  scale — one bf16 ulp of cuBLAS accumulation order. The exactness was
  a property of the workstation's heuristics, not of the code.
* the float32 activation gate is deliberate MLX promotion parity, not
  an accident: layer 3 is the first BailingMLA, whose RoPE path
  computes in float32 and concatenates q_nope.to(q_pe.dtype) with a
  float32 q_pe, so the residual stream is float32 from there on and
  only layers 0-2 (KDA) present bf16 to their dense linears. Reaching
  the rest of the dense 26% is a numerics-parity decision for review,
  not a kernel increment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
…ated

Bring-up plan complete. Everything measured on the target with the
adopted profile (shipped default + EDGE0_MEMORY_BUDGET=auto
EDGE0_CACHE_SLOTS=512 EDGE0_PREDICT_PREFETCH=1), desktop resident.

* Repeat measurement, 3 launches x 2 runs: decode 0.744-0.786 tok/s,
  mean 0.772, sd 0.016 (n=6). Descriptive statistics of a small sample.
* Correctness repeat: same-device reference check of the final profile
  PASSES the registered bound (32/32 tokens, max |dlogit| 0.777 vs 1.0);
  all 30 sustained requests produced byte-identical tokens.
* Sustained 25 min / 30 requests / 960 tokens: mean 0.817 tok/s, sd
  0.028, quartiles 0.791/0.834/0.816/0.824 — no degradation. Junction
  54.5-61.8 C, NO throttling. RSS moved 5.45-5.70 GiB rising AND
  falling (mmapped expert pages reclaimed and re-faulted, not a leak);
  CUDA peaks flat at 1.57 GiB.
* Swap marked explicitly: the run DID swap, +1165 MB in the first fifth
  then flat (+24/0/+18/-2), no throughput loss — a one-time eviction of
  the desktop's idle pages, not the working set. Per Gate C no
  swap-free resident-memory claim is made for this board.
* Envelope: 3292-token prompt validated (prefill 42.4-42.8 tok/s, CUDA
  peak 3.14 GiB, no OOM); generation validated to 256 tokens
  (0.807/0.820 tok/s, no decay).

kv_bytes_per_token for this tier: 1,100,000 -> 600,000. The old value
was the README's Apple/MLX figure; measured here across 128..2048-token
prefills the marginal cost is 561,320 B/token (548 KiB), about half.
Growth is sublinear below the top of the range because most layers of
this hybrid are KDA with fixed-size state and only the MLA layers' KV
grows with context, so the top-of-range slope is the right constant to
extrapolate from. Over-reserving was safe but rejected contexts that
fit: the declared context the budget admits goes ~2560 -> 4096, which
the 3292-token run then validated in practice.

The claim, qualified to what was tested: edge0-8b runs on a physical
Jetson Orin Nano 8 GB at 0.77-0.82 tok/s decode, 42 tok/s prefill at
3.3k tokens, +42% over the Task 3 baseline, within a budget that
rejects what does not fit before loading. Single request, greedy
sampling, <= 4096 declared context, <= 256 validated generation, 25W.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HJKa3V4H3cR9mwKd9qSEbe
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant