We are looking for a Senior Site Reliability Engineer to work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis.
Responsibilities
- Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooks
- Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality
- Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms
- Apply a performance engineering mindset through load testing, capacity analysis, and latency profiling
- Automate operational toil through scripting and infrastructure-as-code
- Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort
- Lead incident response practices including on-call readiness and blameless post-mortems
- Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements
Requirements
- 3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems
- Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events
- Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews
- Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers
- Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines
- Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations
- A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams
- Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations
- English Level: B2+ (Upper-Intermediate) or higher
Nice to have
- Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively
- Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks
- Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation
- Experience with Datadog or similar enterprise observability platforms
- Background in evangelizing best practices and setting standards across engineering teams
- Exposure to programmatic advertising or adtech platforms