Principal Cloud Incident Manager
Seattle, Washington, United States
Drives resolution of highly complex customer-impacting incidents (e.g. Sev1) by running cross-team incident calls, identifying scope and symptoms of issue, pulling in correct resources, appropriately delegating tasks across the global team, and communicating with leadership.Leverages communication skills and business judgement to manage engineers and technical teams triaging large-scale incidents for key customers.Collaborates with key stakeholders during incident post-mortems to define key learnings and deploy preventative solutions.Consistently drives operational improvements to incident management processes, including: Rebuilt companywide Sev1 response process to improve efficiency and information sharing across teams; new process includes updated communication protocols and meeting flow as well as splash pages and timeboxing.Proactively wrote and deployed Python script to automate a time-consuming component of the incident reporting process, reducing time needed from 45 minutes to 4.5 seconds.Leads weekly discussion with cross-company leadership about key incidents and lessons learned, condensing highly technical situations into easy-to-understand briefs.