Skip to content

edge0-8b: ~7-13s to first token on M1 Pro 16 GB, unchanged by warm_willneed / full_layer_prefill #110

Description

@Hamaad616

Environment

  • MacBook Pro (MacBookPro18,1), Apple M1 Pro, 16 GB
  • macOS 26 (build 26A428)
  • Python 3.14.7 (venv from pip install -e '.[dev,fetch]'; not tested on 3.12)
  • edge0 0.1.0, Edge0/Edge0-8B-A1B-preview weights (LoRA and prerouter files present; logs show [lora] applied=153 not_found=0 and prerouter installed: 16 heads, start=7, K=8)
  • Model on the internal SSD, no other heavy apps running, one edge0 process only

What I see
Decode speed is reasonable, but every request has a long delay before the first token, even for a tiny prompt.

Server (edge0 serve, warm process), streaming request with a 27-token prompt. Timestamps on the SSE stream show the first chunk arriving about 10.8s after the server stamped the request, then chunks every ~0.1-0.3s.

Non-streaming timings against the same server:

Tokens generated | Total time -- | -- 12 | 10.5s 19 | 14.7s 200 | 25.7s

The engine reports built in 0.4-0.5s, so it isn't model load time.

*I'm not fully certain the edit applied on this run, since timings are nearly identical to the previous config.

What I checked

  • edge0 models shows prefill_full=0 hot=0 for edge0-8b. As I read the docs, prefill_full_layers=0 means "every layer", and prod_k8() sets full_layer_prefill=True.
  • warm_willneed defaults to False and is only enabled in staged_k4 (35B), so I tried it on the 8B preset. No meaningful change.
  • GPU sampling during a request (powermetrics): GPU active residency 42-67%, ~0.8-1.4 W, i.e. mostly waiting rather than computing.
  • No competing processes. The edge0 server's RSS was ~1.9 GB.

Questions

  1. Is a fixed delay of this size per request expected on 16 GB machines, given the README benchmarks are from a 24 GB M4 Pro?
  2. What is happening in that window? From the docs I suspected whole-layer prefill through the page cache, but toggling the options above made no visible difference.
  3. Is there a supported way to keep more of the model resident between requests, or a setting I should be using on 16 GB?

Happy to run any diagnostics you suggest.


Some notes on the draft before you post it:

  • Check the last table row. I flagged it because the timings looked identical to the previous run. Re-run with full_layer_prefill=False after confirming line 54, and update those numbers (or drop that row) so the report doesn't contain something you can't stand behind.
  • Python 3.14. It's worth trying 3.12 first, since the README recommends it and maintainers may ask. If it makes a difference, that changes the report entirely.
  • Tone check on point 2 of "What I checked". I hedged on what prefill_full_layers=0 means because I'm reading a docstring, not the code. If a maintainer corrects it, that's useful information rather than an error.
  • Consider adding your sysctl output and the exact commands. The curl loop with timestamps from earlier is a good reproducible test.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions