Senior Manager, Site Reliability Engineering
CurrentDrove cross-functional alignment by applying Google Site Reliability Engineering (SRE) principles, including service level objectives (SLOs) and error budget policies. Facilitated agreement among product managers, user representatives, and engineers on acceptable availability levels and predefined remediation actions when targets were not met.Collaborated with leadership and engineering teams to turn around Service Level Agreement (SLA) performance, leading a team of seven Site Reliability Engineers (six of whom I recruited and hired). Under my leadership, the organization met SLAs for six consecutive months following prior missed goals, mitigating contract churn risk and safeguarding revenue.Partnered with finance, vendors, and engineering stakeholders to prevent a significant cost increase with DataDog, achieving a 95%+ cost reduction through detailed analysis, strategic implementation, and effective negotiation.Standardized downtime incident management processes by working across departments to improve communication, consistency, and response efficiency during outages.Led a company-wide migration effort, coordinating across engineering, operations, and customer success teams to rewrite and migrate over 150 customers to a proprietary Expel proxy system, successfully transitioning from CentOS to CoreOS with minimal disruption.