Site Reliability Engineer
Current- Managed and maintained multiple Kubernetes clusters with over hundreds of nodes- Implemented robust alerting and monitoring systems using Envoy, providing real-time insights into egress, ingress, and internal gRPC traffic across 40+ microservices.- Led blameless postmortem meetings to analyze incidents, identify root causes, and implement proactive measures for sustainable incident response.- Developed automation scripts (Ansible, Python) to streamline infrastructure provisioning, configuration management, and deployments, resulting in a 30% decrease in manual intervention.- Implemented machine-to-machine (M2M) authentication and authorization via Keycloak, enhancing security for automated processes.- Collaborated with developers and product owners to define service level objectives (SLOs) for system operations, ensuring consistent performance and availability for our services.- Optimized CI/CD pipelines using caching and multi-stage Docker builds, achieving a 70% reduction in average build times.