Senior Site Reliability Engineer
CurrentResponsibilities:Bake in reliability into architecture, implement adequate automation and controls to monitor the scale and maintain Infrastructure, Network, Storage, API and microservices, in Cloud.A strong ability to design and execute cutting-edge System Testing strategies ( performance/load tests, capacity tests).Working with logging, monitoring, and observability tools (e.g., AWS Cloud Watch, SNS, SES, Lambda, Prometheus, Grafana, etc.).Working in Python and Ansible.learn, research and improve the availability in automationEnsure the availability, performance and scalability of applications in respect of proven design and architecture best practices.Design and execute scalability strategies that ensure the scalability and the elasticity of the infrastructure( network, compute, storage).Own and operate high-volume (transaction/data) environments on the cloud.Extensive experience in blue-green and canary deployments.Experience in managing and scaling high transaction environments.A 50-50 mix between Software Development and System Administration experience.Define the fundamentals of service level measurement to key operational practices like post-mortems and incident response, and in developing service level indicators (SLIs) and objectives (SLOs) for their critical application modules.Recommend approaches and estimated effort for mitigation of reliability gaps, and create and execute reliability assessments on a time to time basis.Experience managing an engineering team on projects with technical deep-dives into code, networking, operating systems and/or storage.Tools: APM , Ansible, AWS cloud watch, SNS, Lambda, Prometheus, Grafana.