Senior HPC DevOps Engineer

EPAM·Poland·Удалённо·2д. назад

We are seeking a Senior HPC DevOps Engineer to join our Science, Innovation & Labs team, responsible for the scaling, reliability, and automation of our high-performance computing (HPC) and machine learning operations (MLOps) platform.

Responsibilities

  • Guide scientists and data teams in navigating and utilizing the platform user interface (UI) effectively, helping them run self-service workloads without direct infrastructure friction
  • Advise users and manage infrastructure capacity regarding capacity blocks versus on-demand usage, optimizing cost, quotas, and resource availability for heavy workloads
  • Maintain automated pipelines for infrastructure provisioning and platform service deployments
  • Resolve technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotas
  • Collaborate with developer experience teams to improve documentation
  • Collaborate with engineering teams to monitor GPU utilization via tools such as CloudWatch or Prometheus
  • Manage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelines
  • Ensure compute availability through capacity planning and reservation management
  • Deploy containerized environments tuned for HPC and GPU pass-through
  • Deploy and scale HPC workloads on cloud infrastructure utilizing parallel storage and networking solutions

Requirements

  • 5+ years of experience in HPC or DevOps engineering roles
  • Knowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirect
  • Proven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/Inferentia
  • Hands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management
  • Experience in deployment of containerized environments using Apptainer/Singularity, Docker, or Enroot
  • Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for Lustre
  • Hands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks
  • Experience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre)
  • Proficiency in English at a B2+ level

Похожие вакансии

Другие вакансии EPAM