We are seeking a Platform Architect to support a customer that develops and manages several HPC clusters across AWS, CoreWeave, GCP and other providers, operating several thousand GPUs today and scaling 10x. This role is Kubernetes-heavy and requires strong software engineering skills, operating multi-cloud platform infrastructure where misconfigurations or failed upgrades cost thousands of GPU-hours.
Responsibilities
- Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale, including cluster lifecycle, node pool management, networking policy, and stability during rapid growth
- Provision HPC infrastructure through CI/CD across AWS, CoreWeave, GCP and OCI, with more providers coming
- Management of job scheduling to allocate GPU compute across training and inference workloads
- Definition and maintenance of SLIs/SLOs, build monitoring and alerting, and participation in incident response and post-incident reviews
- Development of tooling and automation in production-quality code
- Coordination daily with the Networking, Storage, Security and AI/ML platform teams
Requirements
- 7+ years of experience in infrastructure engineering, cloud platforms or HPC
- Expertise in Kubernetes at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
- Proficiency in Python for production-grade tools, not only scripts
- Proficiency in Terraform for writing and reviewing infrastructure as code daily
- Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
- English proficiency at B2 level or higher
Nice to have
- Familiarity with Amazon Elastic Kubernetes Service, Google Kubernetes Engine, Google Cloud Platform
- Background in High-performance computing (HPC), Slurm, Lustre, Amazon FSx
- Skills in Go, Rust, C++
- Familiarity with CI/CD