About the Position
As a Senior Site Reliability Engineer, you will work at the intersection of software engineering and operations, partnering with application, DevOps, and QA teams. You'll enhance observability, automate operational processes, improve CI/CD workflows, and drive reliability initiatives that enable teams to deliver resilient, high-performing software at scale.
About the Project
The project focuses on building and maintaining a highly available, scalable cloud platform that supports business-critical services. Engineering teams continuously improve system reliability, performance, and operational excellence through automation, observability, and modern software delivery practices.
Responsibilities
- Enhance platform observability by designing and maintaining metrics, alerts, dashboards, and monitoring capabilities that improve system visibility and reduce incident resolution time.
- Build and maintain automation, operational tooling, and monitoring solutions that increase service reliability and uptime.
- Work closely with software development and QA teams to embed reliability best practices into software delivery, release processes, and testing strategies.
- Promote operational excellence by driving preventive measures, facilitating blameless post-incident reviews, and supporting capacity and scalability planning.
- Take part in an on-call rotation, ensuring timely investigation and resolution of production incidents affecting critical services.
Requirements
- Strong hands-on experience with Python, particularly for scripting, automation, and operational tooling.
- Proficiency in at least one of the following programming languages: Java, C++, or Go.
- Solid knowledge of Linux environments, cloud platforms (AWS, GCP, or Azure), and containerized infrastructure using technologies such as Docker, Kubernetes, and Terraform.
- Experience designing and maintaining CI/CD pipelines, working with version control systems, and implementing automated testing practices.
- Practical experience with observability and monitoring platforms (such as Prometheus, Grafana, ELK, Datadog, or similar), including troubleshooting through log and metric analysis.
- Experience identifying and documenting Critical User Journeys and translating them into measurable SLA/SLO objectives that support automation and operational excellence.
- Strong collaboration and communication skills, with the ability to work effectively across multidisciplinary engineering teams, especially during critical production events.
- A reliability-first mindset with the belief that system stability is a shared responsibility across engineering teams.
- Familiarity with AI-assisted engineering tools (such as Claude and Codex) and their use within modern software development workflows.
Nice to Have
- Experience developing or maintaining end-to-end and integration tests for distributed or microservices-based systems.
- Knowledge of performance optimization, capacity management, or chaos engineering practices.
- Experience contributing to internal developer platforms, automation tools, or reliability engineering initiatives.
- Understanding of production security, compliance requirements, or change management processes.
- Relevant industry certifications.