Production Engineering Manager
CurrentTeam Scope* Part of Meta’s Infra/Network group* Operate software that runs on all of Meta’s fleet of servers to manage network tuning & class of service marking* Own flow-based network observability across data center, backbone, and edge networks* Includes daemons that run on all fleet and edge machines, BPF programs that run in kernel space, and services running in flexible computeTeam Role* Production Engineering team is embedded with Software Engineers* PE team size is 7-11 people* Use software engineering to ensure that software operates in the fleet reliably and safely* We do coding, design, rollouts, data analysis, dashboards, alert infrastructure, investigations, disaster recovery testing, and cross-team integrationTeam Impact* Measured in hundreds of millions of dollars in infrastructure savings* Reduced in-data center congestion globally* Drove large backbone network customers to an SLA model* Ship reliable software on a regular cadencePersonal Impact* Personally driven multiple initiatives to improve reliability of services and oncall health* When required, played the part of a tech lead to set high-level priorities* Changed the direction of multiple partner team leads to look at possible step changes in how we can approach reliability problems and scaled programs* Often act as the first-person-on-the-ground with a new partner by doing initial outreach and conversations to set context & ground rulesPeople* Managed multiple people through their career growth from junior engineers to tech leads* Hired internal and external team members* Managed for performanceOrganizational Contributions* Owned PE training program at the company & ran hundreds of people through it* Top 10 interviewer for PE* Onboarded multiple Software Engineering partnersOdds and Ends* Been told I'm the only person at Meta who has moved from a management role outside of engineering directly into a Production Engineering Management role