We are seeking a hands-on Senior Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while strengthening reliability, observability, and operational readiness. You will partner closely with engineering stakeholders, contribute during on-call efforts, and drive measurable SLO improvements.
Responsibilities
- Provide on-call support for Java backend services during business hours
- Troubleshoot complex distributed system issues using logs and telemetry
- Identify root causes and drive incident resolution through actionable changes
- Prepare and deploy patches to address cloud infrastructure issues
- Define and improve service metrics and dashboards to assess platform health
- Improve reliability and observability posture for key services
- Create and refine runbooks to standardize operational response
- Track and improve SLOs through repeatable processes
- Submit code changes that improve SLOs when errors occur
- Communicate operational issues clearly and concisely in writing during incidents
Requirements
- 3+ years of SRE/DevOps experience supporting production services
- Strong on-call support experience for backend service ecosystems
- Proven incident response skills using logs and telemetry to find root causes
- Hands-on Amazon Web Services experience
- Solid Amazon DynamoDB experience
- Solid Amazon ElastiCache experience
- Strong Git skills for contributing and reviewing code changes
- Working Gradle knowledge in Java service environments
- Strong observability and troubleshooting skills in distributed systems
- Clear written communication skills for live incident updates
- Fast learning ability to absorb information quickly and apply it under pressure
- English proficiency: B2 Upper-Intermediate
Nice to have
- Kubernetes experience
- Terraform experience
- Grafana dashboarding experience
- New Relic monitoring experience
- Apache Kafka experience