Site Reliability Engineering Manager
CurrentLead a cluster of multiple teams focused on:- Managing high availability and performance throughout the most critical services of the company platform, with a large-scale system design.- Engage in service capacity planning, forecasting and system tuning for periods of high transactional throughput to maintain platform reliability, especially during sale seasons.- Build automation and tooling to prevent problem recurrence, improve MTTR and remove toil.- Incident response and monitoring activities, working following SLAs, SLOs and error rate budgets established for our business flows.- Promote SRE culture with development, architecture and product teams.