View Portfolio: https://akshithmacharla.vercel.app/
Highlights
- Pro
Pinned Loading
-
Parallel-BPE-Tokenizer
Parallel-BPE-Tokenizer PublicHigh-performance GPT-2-style BPE tokenizer in C++ with parallel batch encoding, thread-pool execution, thread-local caching, and benchmark-driven comparison against tiktoken and GPT2TokenizerFast
C++ 2
-
CUDAkernels
CUDAkernels PublicCUDA kernels implemented from scratch for GPU programming, reductions, shared memory, warp-level primitives, and ML/attention kernels.
Python
-
cutlass
cutlass PublicForked from NVIDIA/cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
C++
-
executorch
executorch PublicForked from pytorch/executorch
On-device AI across mobile, embedded and edge for PyTorch
Python
If the problem persists, check the GitHub status page or contact support.

