Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 12 additions & 2 deletions docs/streaming.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,10 +73,20 @@ With `use_compile`, the staged/exact paths are wrapped in `mx.compile`:
| `load_threads` / `prefetch_threads` | Build / prefetch thread counts | 8 / 4 |
| `full_layer_prefill` / `prefill_full_layers` | Whole-layer prefill loading / number of leading layers | False / 0 |
| `prefill_hot` | hot stack size during prefill | 0 |
| `warm_willneed` | Kernel bulk readahead (`madvise WILLNEED`) over the expert ranges a prefetch/stage is about to touch | False |
| `warm_willneed` | Kernel bulk readahead (`madvise WILLNEED`) over the expert ranges a prefetch/stage is about to touch. Does not enable prefetch/staging or affect plain on-demand / whole-layer loading; see below. | False |
| `use_compile` / `top_k` | compile wrapping / routing top-k override | True / None |

Presets: `staged_k4()` (edge0-35b: staged decode with 4 slots, prefill hot stack 32, on-demand prefill), `prod_k8()` (edge0-8b: the reference deployment profile — staged decode off, E3b whole-layer prefill), and `staged_k8()` (the plain K=8 staged variant). Both tiers share `cache_slots=64`.
Presets: `staged_k4()` (edge0-35b: staged decode with 4 slots, prefill hot stack 32, on-demand prefill), `prod_k8()` (edge0-8b: the reference deployment profile with staged decode for prerouter consumer layers and E3b whole-layer prefill), and `staged_k8()` (the plain K=8 staged variant). Both tiers share `cache_slots=64`.

`warm_willneed` is consulted only by `prefetch()` when there are missing experts
and by `stage_experts()` on staged layers. Turning it on does not enable either
path. Disabling history prefetch and staging removes those automatic decode
paths, but a direct `prefetch(experts)` call can still issue readahead for
missing experts. `prefetch_from_prefill()` calls `prefetch()` only if a prefill
expert set was captured while staging was enabled; with staging disabled from
initialization, it is a no-op. The flag does not warm plain on-demand or
whole-layer loads by itself, so it is not a general first-token-latency switch
(see #110).

The whole-layer prefill is the fastest path **when the checkpoint stays in the page cache** (warm 27-token prefill: 0.24 s vs 0.37 s on-demand on an M4 Pro), and the slowest one when it does not (cold: 5.3 s / 4.06 GiB read vs 0.6-1.1 s / 0.4-0.8 GiB; the on-demand figure varies with how many distinct experts the prompt routes to). `edge0 demo|chat|serve --prefill-ondemand` selects the on-demand path for machines in the second group.

Expand Down
7 changes: 7 additions & 0 deletions python/src/edge0/streaming/options.py
Original file line number Diff line number Diff line change
Expand Up @@ -90,6 +90,13 @@ class LayerOptions:
#: a step otherwise degrades into "cold pages x per-fault latency", with
#: those faults serializing on the VM map lock. Retains no MLX arrays,
#: so it does not displace the page cache.
#: Only consulted by ``prefetch()`` (for missing experts) and
#: ``stage_experts()`` (on staged layers). Does not enable either path
#: or affect plain on-demand / whole-layer loading. Direct
#: ``prefetch(experts)`` calls can still consult it when automatic
#: history prefetch and staging are disabled. ``prefetch_from_prefill()``
#: requires an expert set captured while staging was enabled; with
#: staging disabled from initialization, that helper is a no-op.
warm_willneed: bool = False

# ---- presets ----------------------------------------------------------
Expand Down
Loading