Site Reliability Engineer 2
CurrentSRE with the Office 365 Problem Management team. Participate in a 24x7 on call rotation and coordinate incident management (along with incident managers). Escalation point for vendors and FTEs in Operations and other groups in Exchange and externally. Foster relationships between EXSRT and groups in Exchange and externally (development, test, program management, DCS, GFS, GNS, vendors, etc.). Develop, maintain, and document procedures and methods. Engineer and test automation for post mortem data consolidation, incident management processes, etc. Monitoring system health using SCOM and Keynote tools. Coordinate the incident post mortem process and drive to resolution. Coordinating and executing projects with external and internal teams/vendors. Train incoming On-Call Engineers. Some accomplishments in this role include: Deployed the first datacenter installation (Blue Ridge 2) Installed and configured the original Central Admin environment. Developed the first version of the service wide health check suite. (Now integrated into current version). Developed the Operations SharePoint site. Coordinated several new datacenter deployments in the Outlook.com infrastructure. Completed the first decommissioning project for Exchange Online. Fostered relationships with hardware vendors and datacenter teams to reduce time to deploy and repair equipment. Converted Exchange from Definitive Software Library to Team Foundation Services for software access. Developed the first deployment scorecard. Wrote and tested bridge code for automation of SCOM server installation. Wrote and tested bridge code for automation of standardizes server reboots. Wrote preliminary automation for incident data consolidation (Project SHIELD).