Team Lead
CurrentJoined as a key contributor to the Personal Genome Service compute pipeline, focusing on reliability enhancement and turnaround reduction. Migrated compute pipeline from private datacenter into AWS EC2 and ECS. Contributed to both the customer-facing website and internal APIs, driving the release of new features and third-party integrations. Broadening the team's purview to encompass HPC services running in AWS EMR within the Hadoop ecosystem alongside AWS Batch pipelines processing large datasets. * Designed and implemented a cost-saving compute solution for Identity by Decent (IBD) data, slashing expenditures by 80% and resolving operational bottlenecks.* Partnered with research and engineering teams to deliver refined ancestry composition results and support the launch of new Recent Ancestor Location regions and the next-generation Recent Ancestor Locations compute service providing detailed results to those with African American and with Native American ancestry.* Modernized service deployments, transitioning from AMI on EC2 to Docker on ECS, and seamlessly managed the migration of the team’s services and data.* Improved and automated genetic data quality control and access system, eliminating manual processes and streamlining operational efficiency.* Stood up a high-capacity asynchronous Celery compute cluster in AWS, processing over 12 million tasks daily which amplified the functionality of the customer website and optimized batch compute processes.* Implemented an AWS SWF orchestration layer for to ensure reliable task processing, complemented by custom monitoring tools for operational support.* Created pipelines for the phasing and imputation of genetic data and for the computation of derived data for ancestry reports.* Simplified the deployment of Docker-based web applications and services to AWS ECS, streamlining the deployment process with simple configuration files.