Service Reliability Manager
Current• Part of the Major Incident team that senior business & tech stakeholders during the incidents, driving/owning activities the next day(s) until we are safe.• Facilitating & chairing techwide Postmortems where we deep dive across existing processes and tech solutions to identify areas needing improvement, and finally closing out in a business report.• Owning processes like Disaster Recovery, Availability Management, Major Incident Management, producing impactful operational Management info. • Making sure engineering staff are trained in our global operational processes for OnCall.• Understanding and creating useful Management Information (MI) that helps us build a picture of how reliable we are, and being the “point of truth” for Incident Impact understanding.• Using that MI to influence behavior and use it to spot trends and call out areas in need of extra love and effort to reduce risk.