We are looking for a Site Reliability Engineer to strengthen reliability, observability, and platform operations across cloud and Kubernetes environments. In this role, you will improve service health through automation, infrastructure as code and CI/CD practices. Apply now to help keep critical systems stable and scalable!
Responsibilities
- Maintain service reliability by handling L2 operations and incident response
- Operate and administer Kubernetes clusters to ensure stability and performance
- Build and improve CI/CD workflows using Azure DevOps and Azure Pipelines
- Automate operational tasks using scripting to reduce manual effort
- Define and maintain infrastructure as code using Terraform and Ansible
- Implement and refine observability using MELT signals to detect and resolve issues faster
- Coordinate problem resolution by analyzing root causes and proposing corrective actions
- Support secure and resilient cloud operations on Microsoft Azure
- Document operational procedures and share knowledge to improve support readiness
Requirements
- 2+ years of site reliability engineering or DevOps experience
- Kubernetes administration experience supporting production workloads
- Azure DevOps and Azure Pipelines experience delivering CI/CD workflows
- Infrastructure as Code expertise with Terraform and Ansible
- Proficiency in scripting languages for automation tasks
- Strong troubleshooting skills across metrics, events, logs, and traces (MELT)
- Strong understanding of observability concepts and tools
- Azure fundamentals knowledge with AZ-900 or AZ-104 certification (or higher)
- Good communication skills for cross-team incident coordination
- English proficiency: B1+ level or higher
Nice to have
- Argo CD administration or implementation experience
- Experience with Apache Cassandra cluster or timeseries/NoSQL operations
- Familiarity with Grafana and Elastic Cloud or Elastic Stack
- HashiCorp Vault experience
- Knowledge of programming languages, such as Python, Angular, or Go