-
Notifications
You must be signed in to change notification settings - Fork 232
Pull requests: NVIDIA/cudnn-frontend
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
BSA: add Sage FP8 forward support for Blackwell
#475
opened Aug 4, 2026 by
jiayus-nvidia
Contributor
Loading…
test_mhas_v2: remove with_rope from randomized tests
#474
opened Aug 4, 2026 by
vedaanta
Collaborator
Loading…
OSS SDPA prefill: restrict the SM100 engine to sm_100 and skip samples elsewhere
#473
opened Aug 4, 2026 by
vedaanta
Collaborator
Loading…
Fix ragged SDPA backward workspace under-allocation for non-token-major stats layouts
#462
opened Jul 31, 2026 by
YangXu1990uiuc
Collaborator
Loading…
Support FP32 output and dynamic M in row-scaled FP4 grouped GEMM
cat-enhancements
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
orig-nv-eng
Reported or requested by NVIDIA engineering.
#461
opened Jul 31, 2026 by
zianglih
Contributor
Loading…
CSA compressor: review-response fixups for the ratio=128 kernels (follow-up to #427)
#452
opened Jul 30, 2026 by
zkyue
Contributor
Loading…
Add MoE + expert-parallel (MoeEp) Python API with MegaMoE CuTe DSL backend
#448
opened Jul 29, 2026 by
mhoqueanik
Loading…
2 of 6 tasks
Add FFT causal conv1d frontend bindings
cat-enhancements
mod-backend
cuDNN backend API, graph execution, descriptors, engines, or backend integration.
orig-nv-eng
Reported or requested by NVIDIA engineering.
DSA indexer forward: add a lean SM100 fast path for the head_dim=128 / qhead_per_kv_head=64 regime (up to 1.46x)
#416
opened Jul 21, 2026 by
zkyue
Contributor
Loading…
Reject FP8/MXFP8 SDPA forward combinations exposed to a cuDNN 9.24 split-KV bug
#414
opened Jul 20, 2026 by
vedaanta
Collaborator
Loading…
Add version-adaptive nvvm.atomicrmw wrapper for cutlass-dsl 4.5.x compat
#411
opened Jul 20, 2026 by
Anerudhan
Collaborator
Loading…
DSA: fix 576-wide d_sink, test/document top-k index semantics, B300 benchmarks
#405
opened Jul 17, 2026 by
vedaanta
Collaborator
Loading…
Support installing python bindings via cmake --install (fixes #149)
cat-ci
CI failures, test flakiness, workflow breakage, or automation issues.
cat-infra
Build, packaging, tooling, dependency, release, or repository maintenance work.
mod-infra
Infrastructure, CI/CD, build systems, packaging, releases, or repo maintenance.
orig-external
Reported or requested by an external user, customer, or community contributor.
Add large-tensor convolution fuzzer
cat-ci
CI failures, test flakiness, workflow breakage, or automation issues.
mod-backend
cuDNN backend API, graph execution, descriptors, engines, or backend integration.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Add FLOOR_MOD (floored modulo) pointwise mode
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-backend
cuDNN backend API, graph execution, descriptors, engines, or backend integration.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Add MoE + expert-parallel Python API surface, PyTorch reference, and tests
#389
opened Jul 14, 2026 by
Anerudhan
Collaborator
Loading…
DSA: Add FP8/MXFP8 and compressed Top-K indexer paths
#370
opened Jul 9, 2026 by
jiayus-nvidia
Contributor
Loading…
[DRAFT] Add framework-agnostic operator APIs with JAX support
#363
opened Jul 8, 2026 by
mgoldfarb-nvidia
•
Draft
benchmark: compute causal SDPA FLOP counts without allocating masks
#347
opened Jul 5, 2026 by
fallintoplace
Contributor
Loading…
cmake: switch python linker flags to target_link_options
#346
opened Jul 5, 2026 by
fallintoplace
Contributor
Loading…
cmake: include GNUInstallDirs before install interface paths
#345
opened Jul 5, 2026 by
fallintoplace
Contributor
Loading…
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.