Site Reliability Engineer - Observability Platform
Current- Leading the development of Criteo’s next-gen Observability platform.- Collaborating with a team of 4 experts across multiple SRE teams.- Guiding users in setting up effective monitoring for critical applications.- Managing 2000+ Prometheus instances for high-impact monitoring.- Skilled in Mesos, Kubernetes, and Baremetal Infrastructure as Code (Chef).- Committed to reducing technical debt and optimizing infrastructure costs.- Expert at managing complex distributed systems and resolving operational challenges.- Handling on-call responsibilities, ensuring 24/7 system reliability.- Proficient in Python, Golang, Ruby, and Bash for automation and system improvement.- Dedicated to enhancing system reliability, tackling tech debt, and driving cost-efficiency.