Environment
- MacBook Pro (MacBookPro18,1), Apple M1 Pro, 16 GB
- macOS 26 (build 26A428)
- Python 3.14.7 (venv from
pip install -e '.[dev,fetch]'; not tested on 3.12)
- edge0 0.1.0,
Edge0/Edge0-8B-A1B-preview weights (LoRA and prerouter files present; logs show [lora] applied=153 not_found=0 and prerouter installed: 16 heads, start=7, K=8)
- Model on the internal SSD, no other heavy apps running, one edge0 process only
What I see
Decode speed is reasonable, but every request has a long delay before the first token, even for a tiny prompt.
Server (edge0 serve, warm process), streaming request with a 27-token prompt. Timestamps on the SSE stream show the first chunk arriving about 10.8s after the server stamped the request, then chunks every ~0.1-0.3s.
Non-streaming timings against the same server:
Tokens generated | Total time
-- | --
12 | 10.5s
19 | 14.7s
200 | 25.7s
The engine reports built in 0.4-0.5s, so it isn't model load time.
*I'm not fully certain the edit applied on this run, since timings are nearly identical to the previous config.
What I checked
edge0 models shows prefill_full=0 hot=0 for edge0-8b. As I read the docs, prefill_full_layers=0 means "every layer", and prod_k8() sets full_layer_prefill=True.
warm_willneed defaults to False and is only enabled in staged_k4 (35B), so I tried it on the 8B preset. No meaningful change.
- GPU sampling during a request (
powermetrics): GPU active residency 42-67%, ~0.8-1.4 W, i.e. mostly waiting rather than computing.
- No competing processes. The edge0 server's RSS was ~1.9 GB.
Questions
- Is a fixed delay of this size per request expected on 16 GB machines, given the README benchmarks are from a 24 GB M4 Pro?
- What is happening in that window? From the docs I suspected whole-layer prefill through the page cache, but toggling the options above made no visible difference.
- Is there a supported way to keep more of the model resident between requests, or a setting I should be using on 16 GB?
Happy to run any diagnostics you suggest.
Some notes on the draft before you post it:
- Check the last table row. I flagged it because the timings looked identical to the previous run. Re-run with
full_layer_prefill=False after confirming line 54, and update those numbers (or drop that row) so the report doesn't contain something you can't stand behind.
- Python 3.14. It's worth trying 3.12 first, since the README recommends it and maintainers may ask. If it makes a difference, that changes the report entirely.
- Tone check on point 2 of "What I checked". I hedged on what
prefill_full_layers=0 means because I'm reading a docstring, not the code. If a maintainer corrects it, that's useful information rather than an error.
- Consider adding your
sysctl output and the exact commands. The curl loop with timestamps from earlier is a good reproducible test.
Environment
pip install -e '.[dev,fetch]'; not tested on 3.12)Edge0/Edge0-8B-A1B-previewweights (LoRA and prerouter files present; logs show[lora] applied=153 not_found=0andprerouter installed: 16 heads, start=7, K=8)What I see
Decode speed is reasonable, but every request has a long delay before the first token, even for a tiny prompt.
Server (
edge0 serve, warm process), streaming request with a 27-token prompt. Timestamps on the SSE stream show the first chunk arriving about 10.8s after the server stamped the request, then chunks every ~0.1-0.3s.Non-streaming timings against the same server:
The engine reports
built in 0.4-0.5s, so it isn't model load time.*I'm not fully certain the edit applied on this run, since timings are nearly identical to the previous config.
What I checked
edge0 modelsshowsprefill_full=0 hot=0for edge0-8b. As I read the docs,prefill_full_layers=0means "every layer", andprod_k8()setsfull_layer_prefill=True.warm_willneeddefaults to False and is only enabled instaged_k4(35B), so I tried it on the 8B preset. No meaningful change.powermetrics): GPU active residency 42-67%, ~0.8-1.4 W, i.e. mostly waiting rather than computing.Questions
Happy to run any diagnostics you suggest.
Some notes on the draft before you post it:
full_layer_prefill=Falseafter confirming line 54, and update those numbers (or drop that row) so the report doesn't contain something you can't stand behind.prefill_full_layers=0means because I'm reading a docstring, not the code. If a maintainer corrects it, that's useful information rather than an error.sysctloutput and the exact commands. The curl loop with timestamps from earlier is a good reproducible test.