Lead Site Reliability Engineer
Current● Hands on experience in planning and deploying cluster upgrades and maintenance end to end ensuring smooth transition via cluster availability and data integrity.● Full-stack troubleshooting skills across network, application and distributed services layers.● Support services from Dev till Production via infrastructure design, software platform development, stress testing, capacity planning, performance tuning and overall system health etc. Maintain services on clusters and ensure service availability.● Hands on experience in system administration skills, including automation and orchestration of Linux/Windows using Chef, Puppet, Ansible and containers (Docker, Kubernetes, etc.)● Working extensively on IAC technologies like terraform and used it to automate all infrastructure needs in AWS and Azure.● Configured application monitoring and alerts for on-prep and cloud applications using Prometheus with Grafana and Dynatrace.● Responsible for setting KPIs and SLOs.● Proficiency with continuous integration and continuous delivery tooling and practices likeJenkins.● Worked on provisioning and maintaining Docker containers using Kubernetes.● Expert in handing Big data applications support and deploy upgrades. Patching and upgrades with backup and restore capabilities.● Incident/escalation management and reporting with RCA.● Collaborating and resolving platform/environment issues with various network, database andapplication teams.● Worked extensively on HDP and HDF stack including spark, kafka, nifi, hive,Ranger, Ranger KMS, Zeppelin, Knox, Atlas, HDFS, HBASE, Zookeeper,Yarn, configuring highavailability, load balancing, Solr and others.● Experience in Support and maintenance of data Lake, Data Hub, HD Insights, Hadoop,Cloudera, MongoDB, MySQL.● Knowledge of programming and scripting languages such as PowerShell, Bash,SQL, Java, Python● Hands on experience on Kerberos, Ranger and Knox respectively.