We are seeking a Senior Infrastructure Engineer with deep experience in cloud infrastructure, Linux internals, large-scale distributed systems, workload isolation, and a strong security mindset to join our team and help design, build, and operate a global, resilient platform.
Responsibilities
- Design, build, and maintain a global, distributed, and resilient cloud infrastructure
- Collaborate with infrastructure and product engineering teams to plan and deliver complex platform initiatives
- Participate in architecture reviews, incident response, and performance analysis to ensure system reliability
- Manage and provision AWS infrastructure using Terraform and Kubernetes
- Write and maintain Kubernetes manifests and deployment configurations for critical workloads, including pod security contexts, anti-affinity rules, network policies, autoscaling, and health probes
- Drive Production Readiness Reviews (PRR) for all new services, covering security, HA, performance, and observability gates
- Design and operate multi-layer workload isolation using Linux kernel primitives: namespaces (pid, net, mnt, user, uts, ipc), cgroups, seccomp profiles, and capabilities, as the baseline security boundary
- Evaluate and operate gVisor and Firecracker for workloads requiring hard tenant boundaries and near-native performance
- Design and maintain Grafana dashboards
- Manage Prometheus and VictoriaMetrics pipelines; define and tune P1/P2/P3 alert thresholds with runbooks
- Contribute to and extend the internal k6-based load testing framework (load-testing-framework / library/k6/webhooks)
- Design load scenarios using constant-arrival-rate profiles; instrument custom metrics (job_succeeded_count, job_failed_count) tagged by testid for Grafana correlation; stream test metrics to Prometheus via remote write; generate and publish HTML reports to file storage after each run
Requirements
- 5+ years of experience in infrastructure, SRE, or platform engineering roles
- Expertise in distributed systems and cloud-native architectures
- Understanding of Linux internals: namespaces, cgroups, seccomp, capabilities, and system-level performance tuning
- Experience operating infrastructure on AWS at scale
- Proficiency in Terraform and Kubernetes, including security hardening of manifests
- Experience designing and running load tests (k6, Gatling, Locust, or similar)
- Understanding of network security and cloud security best practices
- Excellent analytical, troubleshooting, and communication skills
- Proficiency in English at a B2+ level
Nice to have
- Skills in Golang/Python for building internal tooling
- Familiarity with Kafka, Redis, ClickHouse, or PostgreSQL
- Familiarity with observability tools such as Grafana, Prometheus, or VictoriaMetrics
- Knowledge of encryption key hierarchies (CMK/DEK/KMS patterns) and HashiCorp Vault
- Hands-on production experience with gVisor or Firecracker
- Prior work extending or maintaining an internal testing framework
- Contributions to or maintenance of open-source infrastructure projects