Engineering Manager, Gpu Techniques
CurrentLead a global team of AI engineers to optimize Meta’s inference and training stack for LLMs (Llama) and RecSys models using PyTorch’s ahead-of-time compilation and kernel optimizations on NVIDIA and AMD GPUs.・Launched 26 RecSys models, resulting in annual cost savings of over $100M and contributing to billions Ads revenue.・Expanded team scope to LLMs by developing and integrating Flash Attention 3 and Flash Decoding into production, accelerating Llama inference by up to 4x and enabling long-context inference.・Transformed team identity through scaling AITemplate inference framework and integration of PyTorch 2.0.・Fostered strong collaboration on open-source GPU libraries within/across companies, such as the Meta Fundamental AI Research (FAIR) xformers, OpenAI Triton, NVIDIA CUTLASS, and AMD Composable Kernel.・Managed the diverse team across three time zones.