Site Reliability Engineer
• Identify and provide critical data analysis of systemic issues impacting business critical platforms and systems. • Create metrics and alerts in New Relic, Splunk and Zabbix to monitor highly available business platforms.• Define and advise improvements for the IT infrastructure in the production environment –revolving around availability, capacity, and scaling within AWS cloud architecture.• Provide second level escalation guidance for the OCC and Engineering on high priority incidents impacting business critical systems. • Support the IT Problem Management Process through the use of the ITIL Problem Management framework by discerning, researching and reporting on IT problems.• Drive Root Cause Analysis (RCA) through interviews with subject matter experts (SMEs), data collection and analysis and personal knowledge of business systems and processes to create meaningful corrective actions. • Create problem documentation through RCA’s and problem synopsis for senior management and vendors. • Collaborate closely with Engineering and Infrastructure teams to drive and deliver highly available and durable solutions, including corrective actions generated from RCA’s. • Conduct problem management meetings to help align support resources, diagnose problems and support communication efforts