Site Reliability Engineer
Current- Successfully managing numerous AWS accounts for both production and non-production environments. Key responsibilities encompassed automation, infrastructure buildout, seamless integration, and cost optimization.- Led the deployment and management of Redpanda cluster on EKS, ensuring seamless integration and scalability of Kafka-based messaging systems.- Employing Infrastructure as Code (IAC) practices with Terraform to create standardized AWS infrastructure, effectively facilitating the creation of both non-production and production environments.- Worked on a Centralized inspection architecture, terraforming the code and leveraging AWS Gateway Load Balancer and AWS Transit Gateway to ensure the smooth and secure flow of network traffic, enhancing the overall network performance and security.- Automated the backup of AWS SSM parameters to different regions using Python, ensuring data redundancy and disaster recovery capabilities, while streamlining management and enhancing data availability.- Designed and implemented Python-based backup and restore scripts for Datadog's critical monitoring and dashboard configurations, enhancing operational resilience through custom automation.- Proficiently handling Kubernetes configurations using Helm, including the creation of reproducible builds for Kubernetes applications.