I am an AI infrastructure specialist focused on building, scaling, and optimizing machine learning systems. My work primarily revolves around large language models, distributed training, and cloud infrastructure, with a strong emphasis on Google Cloud Platform (GCP) and Kubernetes.
Currently, I am expanding my contributions to open-source projects, particularly in the AI infrastructure, MLOps, and distributed computing ecosystems.
- Architecting scalable AI infrastructure on Google Kubernetes Engine (GKE).
- Optimizing LLM serving and distributed training using frameworks like Ray and vLLM.
- Streamlining cloud resource provisioning and developing infrastructure automation.
- Identifying and resolving networking bottlenecks in high-performance GPU clusters.
- Cloud & Infrastructure: Google Cloud Platform (GCP), Kubernetes, GKE, Docker, Bash/Shell scripting
- AI & Machine Learning: Ray, vLLM, PyTorch, Distributed Training
- Languages: Python, Shell

