Site Reliability Engineer
Current— maintaining infrastructure in AWS Cloud/SberCloud/YandexCloud— migrating data & clients between cloud solutions— databases managing (SQL and noSQL: Mysql, ElasticSearch, Redis)— implementing automation scripts to improve site efficiency and performance (Python, Bash) — configuring new servers, saving the company time and money, using CloudFormation and Terraform— using Ansible for code deployment and writing ansible playbooks and roles,— Kubernetes, docker containers managing— troubleshooting alerts relating to application failures and devices being unreachable.— monitoring site activity and performance to identify potential issues and recommend solutions using Prometheus, Zabbix, and Grafana— coordinating with developers to resolve site issues and implement changes— providing support for site maintenance and updates— investigating and resolving site outages— collaborating with developers to streamline the build and deployment process, making it more efficient and less error-prone