Senior Network Operations Center Engineer
CurrentJob responsibilities:- Continuously monitoring and analysing key system metrics using various monitoring tools and alerting (I.e. Grafana, Looker, Anadot, Datadog, GCP monitoring, etc);- Verify the automated alerts' conditions; Adjusting the alert conditions or creating Jira tickets for the service owners to create a new alert or adjust the threshold for an existing one;- Performing basic steps to mitigate the issue (I.e. service restart, killing faulty pods, adjusting the service… Show more Job responsibilities:- Continuously monitoring and analysing key system metrics using various monitoring tools and alerting (I.e. Grafana, Looker, Anadot, Datadog, GCP monitoring, etc);- Verify the automated alerts' conditions; Adjusting the alert conditions or creating Jira tickets for the service owners to create a new alert or adjust the threshold for an existing one;- Performing basic steps to mitigate the issue (I.e. service restart, killing faulty pods, adjusting the service scaling policy, etc);- Effectively articulate the existing problems and anomalies to the service owners and work with the teams towards the complete issue resolution;- Documenting all incident-related information for further incident review analysis; - Documenting the Action Items that help to prevent such incidents in the future or speed up the product’s recovery and minimise the downtime; Creating Jiras to hold the track of the AI completion and implementation;__________Personal assignments:- Establishing the processes for the newly created NOC team and adjusting them to the company’s needs (I.e. creating a shift checklist for NOC team; establishing the guidelines for the incident handling process, adjusting the change management process, etc);- Creating and updating the Confluence space of the team; - Holding the meetings with the service owners and creating the service-specific runbooks for the NOC team; - Creating the service mapping for all the micro services; describing how the services work, specifying the user scenarios they are affecting and potential impact during the service outage; Show less