diff --git a/docs/streaming.md b/docs/streaming.md index c049da0..fcb4879 100644 --- a/docs/streaming.md +++ b/docs/streaming.md @@ -73,10 +73,20 @@ With `use_compile`, the staged/exact paths are wrapped in `mx.compile`: | `load_threads` / `prefetch_threads` | Build / prefetch thread counts | 8 / 4 | | `full_layer_prefill` / `prefill_full_layers` | Whole-layer prefill loading / number of leading layers | False / 0 | | `prefill_hot` | hot stack size during prefill | 0 | -| `warm_willneed` | Kernel bulk readahead (`madvise WILLNEED`) over the expert ranges a prefetch/stage is about to touch | False | +| `warm_willneed` | Kernel bulk readahead (`madvise WILLNEED`) over the expert ranges a prefetch/stage is about to touch. Does not enable prefetch/staging or affect plain on-demand / whole-layer loading; see below. | False | | `use_compile` / `top_k` | compile wrapping / routing top-k override | True / None | -Presets: `staged_k4()` (edge0-35b: staged decode with 4 slots, prefill hot stack 32, on-demand prefill), `prod_k8()` (edge0-8b: the reference deployment profile — staged decode off, E3b whole-layer prefill), and `staged_k8()` (the plain K=8 staged variant). Both tiers share `cache_slots=64`. +Presets: `staged_k4()` (edge0-35b: staged decode with 4 slots, prefill hot stack 32, on-demand prefill), `prod_k8()` (edge0-8b: the reference deployment profile with staged decode for prerouter consumer layers and E3b whole-layer prefill), and `staged_k8()` (the plain K=8 staged variant). Both tiers share `cache_slots=64`. + +`warm_willneed` is consulted only by `prefetch()` when there are missing experts +and by `stage_experts()` on staged layers. Turning it on does not enable either +path. Disabling history prefetch and staging removes those automatic decode +paths, but a direct `prefetch(experts)` call can still issue readahead for +missing experts. `prefetch_from_prefill()` calls `prefetch()` only if a prefill +expert set was captured while staging was enabled; with staging disabled from +initialization, it is a no-op. The flag does not warm plain on-demand or +whole-layer loads by itself, so it is not a general first-token-latency switch +(see #110). The whole-layer prefill is the fastest path **when the checkpoint stays in the page cache** (warm 27-token prefill: 0.24 s vs 0.37 s on-demand on an M4 Pro), and the slowest one when it does not (cold: 5.3 s / 4.06 GiB read vs 0.6-1.1 s / 0.4-0.8 GiB; the on-demand figure varies with how many distinct experts the prompt routes to). `edge0 demo|chat|serve --prefill-ondemand` selects the on-demand path for machines in the second group. diff --git a/python/src/edge0/streaming/options.py b/python/src/edge0/streaming/options.py index 8a293f8..6b12fe5 100644 --- a/python/src/edge0/streaming/options.py +++ b/python/src/edge0/streaming/options.py @@ -90,6 +90,13 @@ class LayerOptions: #: a step otherwise degrades into "cold pages x per-fault latency", with #: those faults serializing on the VM map lock. Retains no MLX arrays, #: so it does not displace the page cache. + #: Only consulted by ``prefetch()`` (for missing experts) and + #: ``stage_experts()`` (on staged layers). Does not enable either path + #: or affect plain on-demand / whole-layer loading. Direct + #: ``prefetch(experts)`` calls can still consult it when automatic + #: history prefetch and staging are disabled. ``prefetch_from_prefill()`` + #: requires an expert set captured while staging was enabled; with + #: staging disabled from initialization, that helper is a no-op. warm_willneed: bool = False # ---- presets ----------------------------------------------------------