Site Reliability Engineer
CurrentMember of the SRE team.* Design, build and maintain infrastructure through reusable code and tooling (Ansible, Puppet, Bash, Python scripts, Git);* Create solutions to continuously improve scalability, performance and security of our platform (K8s cluster, Helm, docker, ElasticSearch);* Create and maintain logging stack, monitoring solutions and incident management. (Prometheus, Grafana, Elasticsearch, Elastalert, Filebeat, Fluentbit, PagerDuty, Kapacitor, Netpoller, Influxdb)* Interact with developers, platform and support teams to help them improve availability, reliability and resilience of our infrastructure and systems.* Interact with external companies for support on their appliances (Percona, Netapp, PURE, VMware).* Use your analytical and technical skills to help teams debug and fix issues in different scopes, such as Network, OS, Virtualization, Application.* Spread DevOps best practices across the company.