Senior Role-Site Reliability Engineering
Current- Lead and mentor a team of SRE Engineers, providing technical expertise and mentorship to ensure the successful delivery of projects/tasks/daily operations.- Deploy and manage highly scalable, secure, and reliable AWS cloud infrastructures using best practices and industry standards.- Drive the automation of cloud resource management, Kubenrnetes resource management, scaling operation, deployments, monitoring to improve efficiency and reliability.- Develop/maintain robust CI/CD pipelines for automating application deployments, automating cloud resource creation, and automating production alerts for monitoring.- Collaborate with various teams to ensure the architecture and applications are designed with scalability, reliability and cost in mind.- Develop and maintain monitoring, alerting, and logging solutions to proactively identify and address performance issues and outages.- Mentor and provide guidance to the SREs team, fostering a culture of continuous learning and improvement.- Collaborate with software development teams to incorporate DevOps practices into the software development lifecycle.- Stay current with industry trends, emerging technologies, and best practices to drive innovation and improvements in system reliability.