Site Reliability Engineer
Current- Administered containerization technologies (Kubernetes), proprietary cloud microservices, and physical machines to serve on average 1.4 million simultaneous users- Deployed and troubleshot hyperscale data centers, leveraging Linux expertise to maintain 98% uptime SLAs- Provided detailed reporting and incident analysis through metrics (Grafana) and logs (Kibana)- Collaborated with cross-functional teams to troubleshoot, resolve, and prevent incidents affecting critical services- Adapted and implemented forward proxies from another team, handling 1.3 Tb/s of traffic, leveraging Ansible and CI/CD pipelines to automate deployment and streamline management