I am currently leading some of the most exciting large scale distributed End-2-End Artificial-Intelligent(AI)/Deep -Learning(DL) efforts from inception to production (data curation, training, tuning/customization and deployment). Current focus of my team is AI/DL Algorithm development, performance optimization and ML-Ops Automation. We collaborate with the entire eco system of AI engineering/research teams consisting of architecture/performance, framework (PyTorch, JAX), libraries (CUDA, CuDNN, CuBlas, OpenAI Triton, cuGraph, TRT, NCCL) and distributed infrastructure to facilitate the development of state of the art AI/DL systems, software, libraries and models for Nvidia's GPU community. - Large scale self supervised systems (Generative AI with Large Language Models, Vision, Video and Multimodal)- Reinforcement Learning with Human Feedback ( RLHF) for Generative AI - Graph Neural Networks (DGL, PyTorch-Geometric)- MLPerf-Training and MLPerf-HPC