-
Notifications
You must be signed in to change notification settings - Fork 0
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
Ornith-1.5: validate end-to-end support + plan Qwen3.5 vision
enhancementNew feature or requestNew feature or requestStatus: Open.#468 In pekkah/SharpInference;Sampler: implement XTC (eXclude Top Choices)
enhancementNew feature or requestNew feature or requestpriority: lowLow-priority / backlog; nice-to-haveLow-priority / backlog; nice-to-haveStatus: Open.#466 In pekkah/SharpInference;Sampler: implement no-repeat-ngram (hard ban on repeating an n-gram)
enhancementNew feature or requestNew feature or requestpriority: lowLow-priority / backlog; nice-to-haveLow-priority / backlog; nice-to-haveStatus: Open.#465 In pekkah/SharpInference;Sampler: implement DRY (Don't Repeat Yourself) sequence-repetition penalty
enhancementNew feature or requestNew feature or requestpriority: lowLow-priority / backlog; nice-to-haveLow-priority / backlog; nice-to-haveStatus: Open.#464 In pekkah/SharpInference;docs(readme): refresh the benchmark table — accumulated drift (stale numbers + descriptions) found during #440
maintenanceRecurring upkeep / sync with upstreamRecurring upkeep / sync with upstreamStatus: Open.#442 In pekkah/SharpInference;perf(cuda): DSpark GPU draft round runs ~3× the bandwidth floor — small-M (B=7) projection GEMMs → weight-stationary matvec (#428 lever 1)
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#441 In pekkah/SharpInference;- Status: Open.#437 In pekkah/SharpInference;
GPU Lloyd-Max TqAttention writes compressed-region V aggregate in the rotated basis (no inverse WHT)
Status: Open.#435 In pekkah/SharpInference;KVarN follow-ups: int8-TC prefill GEMM, split-tile decode kernel, minors (#180 successor)
enhancementNew feature or requestNew feature or requestStatus: Open.#433 In pekkah/SharpInference;perf(cuda): CUDA-graph-captured k-token spec verify — the ~10.6 ms fixed verify overhead is now DSpark's whole gap (#428 lever 2)
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#430 In pekkah/SharpInference;- Status: Open.#427 In pekkah/SharpInference;
perf(cuda): dense int8-MMQ prefill ~2–2.6× behind llama.cpp (cp.async MMQ + flash attn) — Qwen3-8B / Gemma4
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#409 In pekkah/SharpInference;