Site Reliability Engineer
Current•Collaborating with cross-functional teams to design and implement a new system architecture and infrastructure,resulting in a 30% increase in system performance and a 20% reduction in operational costs.• Developing automation scripts using Python for routine tasks, reducing manual workload and minimizing human errors and Implementing CI/CD pipelines for seamless deployment and rollback procedures .• Performing Administration and Engineering activities on Open-source Hadoop, Open-source Spark, Airflow, and Machine learning platforms running on Open-source Kubernetes clusters.• Leading and participating in the determination of root-causes for Kubernetes Application service failures andsupporting escalation.• Ensuring Kubernetes platform services effectively met performance and SLA requirements.• Leading Infrastructure Operations and Production Support for container technologies, ensuring seamless and reliable performance.•Designing and developing a cloud-native application using AWS services such as ECS, EKS, and Fargate, resulting in a 40% reduction in infrastructure costs and a 50% increase in application scalability.• Utilizing Shell and Python programming for automation, streamlining repetitive DevOps tasks.• Demonstrating expertise in managing and tuning the performance of Hadoop platforms, including HDFS,Yarn, HIVE, and SPARK. .• Collaborating with cross-functional teams to integrate continuous integration and continuous delivery tools like Jenkins, GitLab, and Terraform, streamlining the software release process.• Participating in on-call rotations, responding to incidents and performing root cause analysis to prevent recurrence.