Site Reliability Engineer
• Implemented AWS Resilience Hub for several applications to provide AWS recommendations, improving application resiliency score by 30%.• Deployed AWS Cloudwatch Alerts using Jenkins CI/CD pipeline with Terraform for several applications to improve observability and decrease detection time for major incidents by 75%.• Architected and implemented robust monitoring solutions using Splunk and Dynatrace to measure SLO/SLI, Application metrics, Availability, Response Time, Saturation and… Show more • Implemented AWS Resilience Hub for several applications to provide AWS recommendations, improving application resiliency score by 30%.• Deployed AWS Cloudwatch Alerts using Jenkins CI/CD pipeline with Terraform for several applications to improve observability and decrease detection time for major incidents by 75%.• Architected and implemented robust monitoring solutions using Splunk and Dynatrace to measure SLO/SLI, Application metrics, Availability, Response Time, Saturation and Errors.• Developed a comprehensive disaster recovery plan with reginal failover and failback to reduce data loss scenarios by 70%.• Reviewed AWS infrastructure test plans and scripts for performance load testing to ensure optimal performance, resulting in a 20% improvement in response time. Show less