Site Reliability Engineer
Accomplishments and Duties: • Primary network engineer for mixed-vendor lab environment with network equipment from Aruba, Cisco, Dell, FS, HP, Juniper, Mellanox, and Supermicro (Cumulus). This lab was used to test enterprise customer equipment to certify that it would work with our primary product, Ubuntu Linux. Since we had minimal control over the type of equipment that customers would send in, this required a high degree of adaptability and the ability to learn new systems on the fly. • Created tooling to aggregate and automate discovery of and management login procedures for over 750 internal and external cloud deployments across multiple bastion "jump hosts". This allowed SREs to search for and log into a management environment for any deployment with a single command, with updates and new deployments published automatically via Terraform. This tool is estimated to have saved the company over $230,000 in wasted man-hours per year. • Developed scripts and procedures to automate migration process for PostgreSQL database deployments, which allowed for rapid mass migration/upgrade of internal deployments from an outdated OpenStack cluster to a newer one in another datacenter. This was part of a massive team effort that allowed our team to clear out a closing datacenter on minimal notice in under 3 months while maintaining downtime SLAs, saving the company from having to pay extremely expensive fees to remain in the datacenter for another month. • Maintained over 750 internal and external cloud deployments via an Infrastructure as Code (IaC) model using Terraform, Kubernetes, Docker, Juju, and Mojo. These environments were spread across multiple private OpenStack clusters as well as public AWS, Azure, and Google Cloud infrastructure. • Coordinated work with teams spread across every continent in a follow-the-sun model. Company is fully remote, so this required very good communication and technical documentation skills.