Sr. Site Reliability Engineering Manager
CurrentLed a team of ten SREs to achieve 99.9% uptime for a system comprising 42 microservices and handling 35M transactions per day with an average response time of 200 milliseconds. Hired, onboarded, and mentored the team, fostering a culture of continuous improvement and innovation, while developing and implementing SRE best practices, methodologies, and strategies. Oversaw design, implementation, and maintenance of scalable and resilient infrastructure. Managed incident response processes and led post-mortem analyses to prevent future issues. Collaborated with development teams to ensure reliability, performance, and security of production systems. Optimized system performance, capacity planning, and cost-efficiency of cloud resources. • Migrated 80% of on-premises services to the Google Cloud Platform (GCP) over 18 months, aligning IT infrastructure with business processes to enhance overall efficiency. • Established and enforced rigorous standards for incident management, change management, and capacity planning, driving operational excellence and reliability across the organization. • Partnered with development teams to embed reliability-focused design and testing methodologies into the software development life cycle, driving significant improvements in system uptime and overall quality. • Designed and implemented comprehensive monitoring systems, using Prometheus, Grafana, and Splunk APM, enabling real-time visibility into system performance and proactive issue detection. • Integrated alerting systems PagerDuty and Slack for seamless notification and incident response, resulting in improved overall system reliability. • Led development and deployment of automated OS update solution, using Rundeck, Puppet, Packer, and BASH, automating patch management and reducing Mean Time to Update (MTTU) by 90%, and resulting in a significant reduction in toil.Trained and mentored domestic and global teams, ensuring seamless knowledge sharing and operational continuity.