We are seeking a Senior HPC DevOps Engineer to join our Science, Innovation & Labs team, responsible for the scaling, reliability, and automation of our high-performance computing (HPC) and machine learning operations (MLOps) platform.
Responsibilities
- Guide scientists and data teams in navigating and utilizing the platform user interface (UI) effectively, helping them run self-service workloads without direct infrastructure friction
- Advise users and manage infrastructure capacity regarding capacity blocks versus on-demand usage, optimizing cost, quotas, and resource availability for heavy workloads
- Maintain automated pipelines for infrastructure provisioning and platform service deployments
- Resolve technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotas
- Collaborate with developer experience teams to improve documentation
- Collaborate with engineering teams to monitor GPU utilization via tools such as CloudWatch or Prometheus
- Manage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelines
- Ensure compute availability through capacity planning and reservation management
- Deploy containerized environments tuned for HPC and GPU pass-through
- Deploy and scale HPC workloads on cloud infrastructure utilizing parallel storage and networking solutions
Requirements
- 5+ years of experience in HPC or DevOps engineering roles
- Knowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirect
- Proven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/Inferentia
- Hands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management
- Experience in deployment of containerized environments using Apptainer/Singularity, Docker, or Enroot
- Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for Lustre
- Hands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks
- Experience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre)
- Proficiency in English at a B2+ level