Site Reliability Engineer
Current- Triage, troubleshoot, and fix production problems in every layer of the stack- Automate the process of deployment and management of SRE tools and services- Help contribute to the design, development, and improvement of logging, monitoring, and alerting- Root cause analysis of incidents and production issues- Participate in an on-call rotation supporting production systems- Document tools and services knowledge, processes, and runbooks- Work across teams to provide technical guidance and to push for best practices