-
Notifications
You must be signed in to change notification settings - Fork 201
Pull requests: Neroued/ninfer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix(qwen3.8): wire-format detect nvfp4 artifact profile
#107
opened Aug 28, 2026 by
koloved
Loading…
Eight weight codes per thread in the MoE prefill staging: routed MoE 1.10x, prefill up to +4.6%
#106
opened Aug 28, 2026 by
MichaelDementii
Loading…
The GDN prefill convolution writes q/k/v itself: 32 GB of copies gone per 32K prompt, 128 MiB of workspace freed
#99
opened Aug 27, 2026 by
MichaelDementii
Loading…
build: cache C++ and CUDA compilation in container builds
#97
opened Aug 26, 2026 by
DuncanBetts
Loading…
fix(ops): size persistent grids from the active device SM count
#89
opened Aug 25, 2026 by
igorls
Loading…
fix(frontend): keep literal vision tokens out of media binding
#88
opened Aug 24, 2026 by
geoffwatts
Loading…
feat(platform): native Windows (MSVC + CUDA) build for ninfer-serve
#84
opened Aug 22, 2026 by
devan-carlin
Loading…
On-demand vision residency: stream the tower through evicted read-only text weights
#72
opened Aug 21, 2026 by
iamwavecut
•
Draft
feat(serve): bound each image with a Vision-token budget
#61
opened Aug 20, 2026 by
Sociopacific
•
Draft
feat(serve): report prefix-cache hits in the chat-completions usage
#55
opened Aug 19, 2026 by
Sociopacific
Loading…
fix(serve): name the exception that terminates the process
#54
opened Aug 19, 2026 by
Sociopacific
Loading…
accept vLLM-style enable_thinking: top-level field and chat_template_kwargs
#50
opened Aug 18, 2026 by
tiequan12345
Loading…
feat(kv): compressed-KV cache (E8 lattice) for the Blackwell sm_120a path
#35
opened Aug 17, 2026 by
danielfparkernz
Loading…
feat(serve): advertise configured model context in models api
#24
opened Aug 15, 2026 by
SolettaSolaris
Loading…
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.