Major It Incident Manager
CurrentLEADING MAJOR IT INCIDENT (P1/P2) INVESTIGATIONSOne of the UK’s Nuclear Weapons Facilities, owned by the Ministry of Defence and operated on its behalf by the UK’s Atomic Weapons Authority, a consortium of organisations: including Jacobs Engineering, Lockhead Martin and the Serco Group. A Contractor role, with HP FDS, reporting direct to the IT Ops Bridge Supervisor and responsible for the identification, diagnosis, management & reporting of all Major IT Incidents. 8000 users on… Show more LEADING MAJOR IT INCIDENT (P1/P2) INVESTIGATIONSOne of the UK’s Nuclear Weapons Facilities, owned by the Ministry of Defence and operated on its behalf by the UK’s Atomic Weapons Authority, a consortium of organisations: including Jacobs Engineering, Lockhead Martin and the Serco Group. A Contractor role, with HP FDS, reporting direct to the IT Ops Bridge Supervisor and responsible for the identification, diagnosis, management & reporting of all Major IT Incidents. 8000 users on a very high security and sensitive site.• Chaired, reported and resolved over 100 major and high priority incidents of service failures. Immediately established potential impact, lost business time and reputation costs; and developed action/communication plans with technical and service teams. • Restored a business intelligence dashboard to full operational service, following a failure costing £28k per hour that breached the service level agreement after 4 hours. Chaired the incident meeting that called in specialist external technical support.• Returned the essential IT services, of a nuclear licensed site, to full operational status within 45 minutes of power being restored, following an unscheduled power outage. Prioritised the essential security, emergency services and critical IT service systems.• Prevented major server and/or system outages, caused by coolant overheating and fan failure, by introducing more effective processes and procedures for monitoring and reporting of 260 data centre rack temperatures, that enabled faster responses.• Improved technical reaction and system restoration time; and communication levels to business users, following major P1 and P2 incidents. Introduced a more efficient reporting process that ensured a more calm, structured and consistent approach. Show less