Lead Site Reliability Engineer
CurrentAs the Lead Site Reliability Engineer, I am responsible for building an SRE team and culture that aligns with corporate goals by continually improving system reliability, resilience, and efficiency.Sheetz’s IT landscape is comprised of tens of thousands of endpoints consisting of traditional IT assets (virtualized or containerized applications running on virtual machines or k8s clusters, running on premises, in the public cloud, or in a private cloud), embedded solutions (fuel pumps, self-checkout devices, car washes, point of sales devices, etc.), running Windows and Linux, supporting multiple unique lines of business.To date, I have:• Enabled executive leadership to see the impact of outages on revenue and customer experience• Enabled executive leadership to make informed new capability-vs-improved reliability trade-off decisions by establishing, measuring, and reporting Service Level Objectives and error budgets• Enabled microservice owners to identify and act on episodic and systemic issues that impact reliability and availability by initiating and overseeing the creation of customized actionable observability dashboards• Improved the organization’s ability to identify and remove incident root causes by implementing six-sigma 5-why activities and project management behaviors• Reduced false alarm alerts by leading the effort to evolve from symptom-based to problem-based alerting• Began laying the ground work for incorporating an AI Ops Platform into business processes by initiating and leading the effort to define logging policies and standards, and by establishing an AI Ops working group comprised of cross-functional stakeholders.• I have defined the Vision, Mission, Strategy, Alignment, and Staffing Model of a future SRE organization, and I continue to work collaboratively across the organization to build consensus and seek approval.