Senior Principal Site Reliability Engineer / Infrastructure Architect
Spearheaded architecture and maintenance of Sling’s production and development infrastructure to enhance system reliability and operational uptime. Lead incident response initiatives to minimize system downtime and define and deploy SLOs and SLAs to maintain and exceed service reliability standards. Serve as the primary technical authority and influence all strategic platform engineering decisions.• Facilitated the transition of legacy infrastructure to Infrastructure as Code (IaC) using Terraform and AWS CloudFormation.• Designed and implemented cloud and data center solutions, including AWS migrations, Kubernetes cluster management, and establishing Direct Connect links to integrate on-premises data centers with AWS cloud.• Designed and implemented cloud and data center solutions including AWS migrations and Kubernetes cluster management which scaled across eight physical data centers.• Drove organization-wide security enhancements by architecting Palo Alto firewall solutions and ensuring adherence to corporate security policies.• Deployed and integrated Harmonic into existing video workflows• Mentored team and developed a skilled workforce adept in SRE, DevOps, and Platform engineering.• Led full organizational and process redesign twice to streamline operations and reduce silos.• Spearheaded end-to-end Kubernetes infrastructure including overseeing application migrations and system integration.• Manage hardware and software assets including ~20K servers across VMware and AWS environments, as well as various custom-built storage solutions.• Provide technical leadership and hands-on training to the site reliability team to enhance team capabilities and improve overall system performance.• Oversaw multiple storage solutions such as custom-built, Pure, and Isilion and network devices including switches, routers, and firewalls to ensure stable and secure operations.