I build and benchmark GPU workloads focused on understanding how hardware, memory, and distributed computing impact modern AI systems.
My work focuses on:
| Focus Area | Technologies |
|---|---|
| GPU Programming | CUDA, CUDA Memory Management, GPU Optimization |
| Distributed Training | PyTorch DDP, NCCL, DistributedSampler |
| Model Scaling | Tensor Parallelism, Transformer Parallelization |
| Deep Learning Systems | PyTorch, Mixed Precision Training |
| Languages | Python, C++, C |
|
Multi-GPU training implementation using PyTorch distributed systems.
|
Explored model parallelism techniques for transformer architectures.
|
CUDA Programming
↓
GPU Memory Optimization
↓
Distributed Training
↓
Efficient AI Infrastructure
Currently exploring:
- GPU performance optimization
- CUDA-based acceleration
- Distributed deep learning systems
- Efficient AI workloads
| Category | Tools |
|---|---|
| GPU Computing | CUDA, CUDA Toolkit, GPU Memory Management |
| AI Frameworks | PyTorch |
| Distributed Systems | DDP, NCCL, Tensor Parallelism |
| Programming | Python, C++, C |
| Development | Linux, Docker, Git |
| Area | Highlights |
|---|---|
| GPU Computing | CUDA memory allocation benchmarks, GPU optimization experiments |
| Distributed Training | PyTorch DDP with NCCL multi-GPU communication |
| Model Parallelism | Tensor Parallel Transformer implementation |
| Computer Vision | GPU accelerated EuroSAT training pipeline |
| Performance Engineering | Runtime analysis, memory profiling, workload benchmarking |


