High Performance Computing (Hpc) Lab Architect
Washington, Dc, Us
• Primary Architect for the CFDLab (UT Austin), Aerolab, and Flight Sciences Labs (NASA/JSC) high-performance computing Linux environments. Led system life-cycle from procurement, through deployment, to operations and retirement for our network, storage, compute, and workstation hardware.– Architected 3 generations of Lustre Filesystems, from 2007 to present. Selected all server components, MDS/OSS width and depth balance, storage subsystem configuration & tuning, deployment, operations, and retirement. Production Lustre experience ranges from v1.6 to v2.14, in a multi-homed InfiniBand/40GbE network environment, including both ldiskfs & ZFS MDT & OST technologies, and active/active (MDT) & active/passive (OST) failover pairing.– Architected/selected 7 generations of Linux clusters, from 2000 to present. Experience dates from very early ‘Beowulf’ 16-node clusters, to (currently) multiple InfiniBand-connected HPE ICE-X/XA systems & federated under a single SLURM interface with a companion Supermicro blade cluster.– Architected a distributed lab environment across the NASA/JSC campus, including network and workstation selection, enabling users’ desktop & software environment to seamlessly match centralized HPC resources, including mounting the same home and Lustre filesystems. This provided a significant efficiency for our use cases, where large files are often required for pre- and post-processing of analyses. Users can develop inputs for PDE analysis directly on the HPC filesystem, and post-process in the same global workspace. Importantly, data transfer for HPC staging was largely eliminated in this setting.