Senior Engineering Manager (Infrastructure /Devops /Site Reliability)
CurrentMy team is Responsible for the Reliability of all Conversational Modeling Engineering Infrastructure . Below are the highlights of my contributions at LivePersonResponsible for Automation of AWS/GCP Cloud Environment Build using IaC (Terraform) Owned Observability solution for GKE using Prometheus and Grafana.Maintenance of GKE cluster and full operations support Creating SLO's and Error budget based on SLA's agreed for the web services and managing error budget to satisfy the SLO requirements . Used automation to reduce toilLed CI/CD Deployment automation for GCP/AWS workloads on Kubernetes using Gitlab and FluxCD .Optimized GPU resources and strategy to manage costs for training and inference of models in GCP for DataScience Org. Reduced costs by 40% .Headed the development and execution of the entire pipeline, encompassing LLM model training using Transformers library in conjunction with PyTorch, deployment, and load testing, ensuring seamless integration and scalability in machine learning workflowsHired and Setup SRE/Devops Team from ScratchPresentations to management on Health of Systems / New Projects Rollout planWorking closely with development, operations, and product teams to align priorities and drive operational excellence.