Infrastructure engineer with 12+ years of experience, including time at Amazon Web Services, now focused on the operational backbone behind large-scale AI/ML systems β Kubernetes, GPU orchestration, and the reliability tooling that keeps model-serving infrastructure observable, debuggable, and recoverable at scale.
- βοΈ 12+ years in cloud infrastructure and DevOps/SRE, including AWS (Cloud Support Engineer, Technical Account Manager)
- π€ AI infrastructure: GPU cluster management, Kubernetes operators for ML workloads, Slurm/batch orchestration
- π§ Go, Python, Terraform, Kubernetes/EKS, AWS Batch
- π§ Agentic AI: built an AI property assistant (chatbot) and automated content generation for PropX8.com using LangChain/LangGraph
- π Recent work:
- PropX8.com β AI-backed real estate platform (Mumbai) with an AI property assistant and agent-driven automated blogging, built at SkywardTech
- gpu-guardian-operator β Kubernetes operator for GPU node health (Xid/ECC errors, thermal throttling), zero external dependencies
- terraform-aws-eks-gpu-nodepool β end-to-end Terraform stack for GPU workloads on EKS (VPC, cluster, spot/on-demand GPU node pool)
- π« Reach me: milind2sisodiya@gmail.com