Site Reliability Engineering Manager | Digital Services
CurrentAs the Site Reliability Engineering (SRE) Manager for Taco Bell’s Digital Services, I oversee the reliability and performance of the brand’s Ecommerce platforms, mobile applications (iOS, Android, Web), and middleware services. I lead a talented team of SRE engineers and reliability engineers, ensuring seamless and highly available digital experiences for millions of Taco Bell customers. Utilizing AWS and serverless technologies, I focus on building scalable, efficient systems and ensuring continuous improvements through automation and best practices.Key Responsibilities: • Manage, mentor, and develop a team of SRE engineers dedicated to supporting and enhancing Taco Bell’s digital platforms. • Lead observability efforts, ensuring our monitoring systems provide critical insights into application health and performance across digital platforms. • Partner with cross-functional teams, both technical and non-technical, to drive operational excellence and manage incident responses, including working with third-party platforms and vendors. • Drive modern SRE practices such as implementing SLIs, SLOs, self-service tools, and reducing operational toil through automation. • Lead Agile ceremonies, ensuring proper requirements gathering, sprint planning, and optimal distribution of work to meet team objectives. • Play an integral role in the innovation and transformation of Taco Bell’s digital services, ensuring a reliable, scalable, and customer-centric experience.Skills & Technologies: • AWS (focus on serverless architecture) • Observability tools & platforms (e.g., CloudWatch, Datadog) • Automation, CI/CD, and infrastructure as code • Incident management and blameless postmortems • Strong understanding of Agile methodologiesIn this role, I not only lead but help shape the future of Taco Bell’s digital infrastructure, driving continuous improvement to meet both customer expectations and business objectives.