We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java backend services ecosystem alongside another SRE and a backend engineering team. You will strengthen reliability, observability, and incident response practices.
Responsibilities
- Provide on-call support for Java backend identity services during business hours
- Troubleshoot complex distributed-system issues using logs and telemetry to identify root causes
- Deploy patches to address issues in cloud infrastructure
- Improve reliability posture for key backend services through tangible changes
- Build metrics and dashboards to quickly assess overall platform health
- Monitor SLOs across services and submit code changes to improve SLOs as errors occur
- Create and refine runbooks for backend services to standardize operations and response
Requirements
- 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems
- 5+ years of experience with Amazon Web Services in production environments
- 3+ years of experience with Amazon DynamoDB and Amazon ElastiCache in production
- Leadership experience guiding on-call practices and operational improvements
- Strong project execution skills delivering reliability and observability enhancements
- Strong troubleshooting skills using logs and telemetry to find root causes
- Strong communication skills for clear and concise written incident updates
- Fast learning ability to absorb information quickly and apply it during on-call
- Upper-Intermediate English proficiency (B2)
Nice to have
- Kubernetes
- Terraform
- Apache Kafka
- Grafana
- New Relic