Senior Site Reliability Engineer - Lead Major Incident Manager
• Ad-Hoc Team Building and Management:o Assembled and led cross-functional teams to address critical incidents, leveraging the expertise of various departments to drive swift resolution.o Fostered a collaborative environment, ensuring team members were aligned on objectives and responsibilities.• Leadership and Communication:o Served as the primary point of contact for major incidents, coordinating response efforts and ensuring clear communication across all levels of the organization.o Effectively communicated with C-suite executives, providing timely updates on incident status, impact, and resolution strategies.o Acted as a liaison between technical teams and business stakeholders, translating technical information into business-relevant terms.• Technical Analysis and Incident Management:o Led the technical analysis and troubleshooting of complex IT incidents, ensuring rapid resolution and minimal impact on business operations.o Utilized monitoring, alerting, and observability tools to quickly triage and remediate critical technology incidents. o Conducted root cause analysis to identify underlying issues and implemented preventive measures.o Managed the incident lifecycle using ITIL best practices, from detection and diagnosis to resolution and closure.• Documentation and Reporting:o Maintained comprehensive documentation of incidents, including detailed incident reports, timelines, and post-incident reviews.o Developed and updated incident response procedures, ensuring they were current and aligned with industry best practices.o Prepared and presented incident reports and metrics to senior leadership, highlighting trends, lessons learned, and areas for improvement.• Process Improvement:o Identified opportunities for automation and tool enhancements to streamline incident management activities.o Led initiatives to implement best practices solutions to improve overall incident response capabilities.