We are looking for a Senior to Staff Infrastructure Engineer to lead the design and evolution of large-scale, multi-GPU compute infrastructure used to train next-generation robotics and AI models. This role sits at the intersection of DevOps, MLOps, and distributed systems — owning architecture, reliability, and performance at scale in a fast-moving, cutting-edge environment.
Essential functions
- Lead architecture and long-term technical direction of multi-GPU, cross-cloud training platforms
- Build and evolve infrastructure-as-code for provisioning, orchestration, and lifecycle management
- Architect and improve CI/CD systems for infrastructure and ML training workflows
- Optimize distributed training workloads — scheduling, resource utilization, observability
- Partner with ML engineers and researchers to enable efficient experimentation and productionization
- Mentor engineers and drive operational excellence across the org
- Document architecture, systems, and key technical decisions
Qualifications
- Production-grade Kubernetes experience (CKA preferred)
- Hands-on Terraform (infrastructure-as-code)
- Kubernetes packaging and release management via Helm
- AWS cloud operations experience
- CI/CD pipeline experience including self-hosted runners (GitHub Actions)
- Prometheus/Grafana monitoring and alerting
- Linux administration, containerization, scripting (Python & Bash)
- Availability for on-call rotation
Would be a plus
- GPU-accelerated Kubernetes clusters (NVIDIA)
- Cluster autoscaling (Karpenter)
- Workflow orchestration (Prefect)
- Gang scheduling, fair-share resource allocation
- High-performance storage (FSx for Lustre, EFS)
We offer
- Opportunity to work on bleeding-edge projects
- Work with a highly motivated and dedicated team
- Competitive salary
- Flexible schedule
- Benefits package - medical insurance, sports
- Corporate social events
- Professional development opportunities
- Well-equipped office
About us
Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI,
and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical
challenges and enable positive business outcomes for enterprise companies undergoing business transformation.
A key differentiator for Grid Dynamics is our 8 years of experience and leadership in
enterprise AI, supported by profound expertise and ongoing investment in
data,
analytics,
cloud & DevOps,
application modernization
and
customer experience.
Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.