Senior Site Reliability Engineer
Current• Led a team that successfully implemented Kafka, KSQL, and Flink, resulting in reduced plant downtime, improved on-time delivery, and $15 million in additional revenue.• Automated Neo4j graph database operations, enabling swift correlation of order, shipment, and external site information.• Improved the availability of critical applications from 99.95% to 99.9999%, meeting service level objectives and enabling prompt responses to AWS outages.• Optimized instance sizes and utilized spot instances, leading to a 60% reduction in resource consumption, 70% reduction in cost, and annual savings of $500,000.• Designed and executed an identity federation solution employing Auth0 and Ping, facilitating secure access to applications via OIDC or SAML protocols.• Designed and developed authorization mechanisms for cross-account and multi-region AWS environments, leveraging OpenID Connect (OIDC) as the identity provider.• Established Datadog APM monitoring for both EKS and serverless environments, implemented dashboards and alerts, and mentored engineers for effective utilization.• Set up Gitlab and Azure DevOps CI/CD pipelines, optimizing deployment times, incorporating security measures, and integrating approval processes, resulting in a decrease in deployment time from 2-3 weeks to 0-1 days.• Improved the efficiency and quality of the development process through automation, security enhancement, CI/CD optimization, and mentoring engineers.