We are seeking a Lead Site Reliability Engineer to embed with the team and drive reliable operations across a broad infrastructure landscape. You will own monitoring, observability, and logging with Dynatrace and Splunk, and help evolve dashboards, alerting, and analysis practices. You will also join the on-call rotation and partner on incident response. Apply now.
Responsibilities
- Monitor and sustain the health, performance, and reliability of the client's applications and services
- Operate and continuously enhance Dynatrace and Splunk, including dashboards, alerting, anomaly detection, and log analysis
- Analyze alerts, determine likely root causes, and deliver actionable recommendations to engineering and incident management teams
- Support incident response activities and participate in the on-call rotation
- Perform Root Cause Analysis (RCA) and drive measurable post-incident improvements
- Define and monitor SLOs, SLAs, and error budgets
- Detect and remediate observability and monitoring gaps across services
- Maintain operational runbooks and supporting documentation
Requirements
- Proven background with 5+ years in Site Reliability Engineering, Production Operations, DevOps, or a closely related role
- Deep expertise in Dynatrace and Splunk, including APM, alerting, dashboards, RUM, synthetic monitoring, service flow analysis, SPL queries, and log analysis
- Hands-on experience with production incident management and on-call support, covering alert triage, incident response, RCA, and post-incident reviews
- Strong troubleshooting skills and RCA capability across distributed applications and services
- Solid understanding of application architecture, service dependencies, integrations, performance analysis, dependency mapping, and bottleneck identification
- Practical experience with AWS services, including CloudWatch, ECS, EC2, ALB, Route53, RDS, and VPC
- Working knowledge of CI/CD pipelines and release validation processes
- English proficiency at B2 (Upper-Intermediate) level or higher
Nice to have
- Familiarity with AI-assisted observability capabilities
- Experience optimizing monitoring and alerting strategies
- Exposure to Infrastructure as Code (Terraform or equivalent)
- Travel/Airline industry experience
- Background supporting modernization and cloud transformation initiatives