Site Reliability Engineer
Current• Led DevOps migration and automation initiatives for system build and deployment processes, establishing infrastructure across Dev, QA, UAT, and Production environments spanning remote, on-premises, and cloud platforms.• Defined and monitored Service-Level Objectives (SLOs) and Service-Level Agreements (SLAs) to ensure system reliability aligns with business goals, achieving a 98% compliance rate with established targets.• Led incident response for critical system outages, performed root cause analysis, and conducted postmortems to implement preventive measures, reducing incident recurrence by 75%.• Automated security compliance checks and vulnerability assessments using AWS Config, CloudTrail, and Security Hub, enhancing security posture and reducing manual audit efforts by 60%.• Developed Terraform code to automate AWS infrastructure setup and deployment across various cloud, server, container, and environment configurations.• Configured AWS resources, including EC2, EBS, security groups, auto-scaling instances, load balancers, regions, AZs, and VPCs.• Wrote Terraform scripts to provision AWS resources like EC2, EBS, security groups, S3 buckets for state files, and DynamoDB for backend files.• Implemented CI/CD for web applications on AWS using services like Code Commit, Code Build, Sonar Cloud, S3, Code Artifact, Code Pipeline, Code Deploy, Elastic Beanstalk, and RDS.• Automated setup of cloud stacks with Ansible for services such as Tomcat, Memcached, RabbitMQ, and MySQL.• Utilized Python and Bash scripting for server-related tasks and interfaced with AWS services using the python-boto module.• Developed Jenkins pipelines to support deployment processes and implemented CD with Jenkins and Ansible.• Created Docker-based CD pipelines with Jenkins and GitHub, automating container builds for new GitHub branches.• Developed scripts using Jenkins, Docker, Maven, and Python for build, deployment, maintenance, and related tasks.