Site Reliability Engineer
CurrentOperational Support & Monitoring• Serve as the primary operational support for large-scale, distributed applications deployed on AWS and GCP, ensuring system reliability and optimal performance.• Monitor system availability and evaluate health metrics, implementing proactive alerts to address potential issues before escalation.• Troubleshoot and resolve production issues across services and hosting stack layers.• Configure and manage Grafana Cloud with Prometheus for metrics monitoring and Loki for log aggregation, enhancing system observability. Infrastructure• Design, build, and maintain scalable infrastructure to support thousands of concurrent users across AWS and GCP environments.• Cloud Migration: Contributed to the migration of critical workloads from AWS to Google Cloud Platform (GCP), including:1. Migrating applications from EC2 to Compute Engine.2. Transitioning containerized workloads from EKS to GKE.3. Migrating relational databases from RDS to Cloud SQL.4. Managing container images in Google Container Registry (GCR). CI/CD and Automation• Developed and implemented CI/CD pipelines using GitHub Actions to enable efficient integration and deployment processes.• Automated deployment tasks and managed configurations with Ansible.