Here are
3 public repositories
matching this topic...
Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode. New: opt-in uncensored mode (runtime abliteration, no new weights).
Updated
Oct 2, 2026
Python
Qwen3.8-Flash-Next 177B MoE (ISTA GSQ-RCO IQ3_XXS) self-hosted on RTX 5090 Laptop (24GB VRAM + 64GB RAM) via Strata: up to 110 tok/s (100+ sustained), 256K full context, vision, MTP speculative decoding, reasoning default xhigh (max tier). GPU+CPU both saturated. 26 measured rounds, 7 quant tiers screened. 中英双语实测实录
Updated
Oct 5, 2026
Python
QSA HiSparse for SGLang: 256K KV offload, CUDA Graph benchmarks, and patches tested on dual RTX 4090 48GB
Updated
Oct 5, 2026
Python
Add this topic to your repo
To associate your repository with the
256k-context
topic, visit your repo's landing page and select "manage topics."
Learn more
You can’t perform that action at this time.