Site Reliability Engineering Manager
CurrentLead a global, fully remote team of 12 engineers across three continents, focusing on production SaaS cloud reliability, capacity, security, observability, and cost optimization.Hands-on involvement in day-to-day operations, including:- Upgrades and production releases- Building tools to improve operational efficiency- Coordinating cross-team projects- Manage infrastructure primarily running on AWS EKS, utilizing a broad range of AWS services and cloud-native technologies.- Actively working on a FedRAMP-compliant sovereign cloud project, ensuring security and regulatory compliance.- Scaled the SaaS service from 3 to over 20 highly-available geo-distributed Kubernetes clusters with over a million transactions a day, driving operational improvements and capacity.- Grew the SRE team from 2 engineers to 12, fostering a culture of collaboration and operational excellence.- Strong focus on open-source tools and cloud-native technologies to ensure scalability and reliability.