Site Reliability Engineer
• Provided key support to the Large Languages Model Team in developing a Python system for integration of LLM based summary the core product. A system that required a large degree of accuracy due to the sensitivity of the data.• Researched and Implemented specific compliance SOC-2 based rules in Terraform based IaC. Ensuring compliance and automating work away from the Ops team.• Created new GitHub Issues templates to enhance deployment requests, reducing process time by 10% and… Show more • Provided key support to the Large Languages Model Team in developing a Python system for integration of LLM based summary the core product. A system that required a large degree of accuracy due to the sensitivity of the data.• Researched and Implemented specific compliance SOC-2 based rules in Terraform based IaC. Ensuring compliance and automating work away from the Ops team.• Created new GitHub Issues templates to enhance deployment requests, reducing process time by 10% and overall making the process more manageable.• Monitored and maintained daily database backups, as well as migrating critical databases during deployments in both production and non-production, ensuring the 98% uptime SLA by performing these tasks during off-peak hours.• Improved monitoring and logging visualizations with Grafana and Prometheus, leading to a 15% reduction in time spent interpreting the graphs.• Created a groundbreaking retrospective meeting that led to a 20% reduction in release process errors, fostering a culture of candor-based blameless improvements to our systems.• Oversaw the deployment and a weekly operations rotation of both production and non-production AWS environments with complex network configurations. Achieving a 97% uptime by doing late night deployments.• Successfully resolved >3 support tickets weekly by implementing Infrastructure as Code (IaC) and declarative configuration management using Terraform, Nix, and NixOS, resulting in a significant reduction of the backlog. Show less