Snr. Systems And Infrastructure Engineer
Current● Oversee the deployment and configuration of AI and HPC systems, including setting up operating systems, drivers, and optimizing server with GPU for maximum performance● Designed and architect hardware and software infrastructure for AI and HPC applications, including selecting and configuring servers, storage, and network components● Diagnosed and resolved hardware and software issues that may arise, ensuring minimal disruption to ongoing AI and HPC operations and conducted performance benchmarking and tuning to maximize the efficiency of AI models and HPC workloads● Linux Plumbing & Kernel Eng: Maintain Linux kernel and core userspace subsystems including submitting patches upstream against latest stable releases, with a focus on networking, and optimizing performance on Linux hosts with dozens to hundreds of CPU cores● Work closely with our internal customers, software developers, and other stakeholders to align kernel development with overall project goals, including customer requirements for GPU-based solutions and other specialized needs.● Leveraged Ansible and Bash scripting to manage over 3000 remote servers from a centralized jump box, maintaining system architecture using Grafana for performance monitoring and continuously monitor system performance and health, using tools and techniques to identify and address issues proactively.● Supervised and managed processes and applications, including Group Policy, ADDS, and client systems across 2900 devices. Performed migration/upgrade from RHEL 7 to RHEL 8 and installed Certbot on Linux open source and managed it● Proficient in configuring and managing storage solutions using logical Volume Manager (LVM) to optimize disk space allocation, create logical volumes, and implement snapshots for data backup and recovery purposes● Configured and optimized DNS servers for efficient resolution and load distribution and implemented DNS security measures such as DNSSEC and DNS firewall to enhance system protection.