Join EPAM as a Senior Data DevOps Engineer and take ownership of building, automating, and evolving reliable, scalable, and secure infrastructure for data-intensive environments.
In this role, you'll work across AWS and Kubernetes environments, automate infrastructure with Terraform and Ansible, build and optimize CI/CD pipelines, and improve the reliability and observability of data platforms. You'll collaborate closely with data engineers, developers, architects, and other stakeholders to solve complex technical challenges and establish efficient DevOps practices.
Responsibilities
- Design, deploy, and maintain scalable and secure AWS infrastructure for data platforms and workloads using Terraform and Infrastructure as Code (IaC) best practices
- Manage and optimize Kubernetes environments, ensuring reliable deployment, scaling, and operation of containerized workloads
- Develop and maintain automation with Terraform and Ansible to streamline infrastructure provisioning, configuration, and operational processes
- Design, maintain, and optimize CI/CD pipelines to enable reliable and automated application, infrastructure, and data platform delivery
- Implement and maintain monitoring, logging, alerting, and observability solutions across AWS, Kubernetes, and data environments
- Troubleshoot and resolve complex infrastructure, deployment, networking, performance, and availability issues
- Develop scripts and automation using Bash and Python to reduce manual effort and improve operational efficiency
- Collaborate with data engineers, developers, architects, security teams, and other stakeholders to design and implement effective infrastructure and DevOps solutions
- Ensure cloud and Kubernetes environments follow best practices for security, scalability, reliability, performance, and maintainability
- Identify opportunities to improve platform reliability, deployment processes, automation, and operational procedures
- Participate in architecture discussions and contribute to technical decisions related to cloud infrastructure, Kubernetes, automation, and DevOps practices
- Support production environments, participate in incident resolution, and proactively address operational risks and recurring issues
Requirements
- 5+ years of practical experience in DevOps, Cloud Engineering, Infrastructure Engineering, or a related field
- 3+ years of hands-on experience with AWS and Infrastructure as Code
- Strong hands-on experience with Terraform for provisioning and managing cloud infrastructure
- Solid experience with Kubernetes and containerized environments
- Practical experience with Ansible for configuration management and infrastructure automation
- Strong experience designing and maintaining CI/CD pipelines using GitHub Actions, GitLab CI/CD, Jenkins, or similar tools
- Good knowledge of core AWS services and cloud architecture principles
- Strong understanding of AWS networking, IAM, security, scalability, and reliability best practices
- Good scripting and automation skills with Bash and Python
- Hands-on experience with monitoring, logging, alerting, and observability solutions
- Strong troubleshooting and incident-resolution skills across cloud, Kubernetes, infrastructure, and CI/CD environments
- Good understanding of DevOps principles, Infrastructure as Code, automation, and CI/CD best practices
- Ability to collaborate effectively with data engineers, developers, architects, and other technical stakeholders
- Upper-Intermediate or higher level of English
Nice to have
- Experience with Amazon EKS and Kubernetes in production environments
- Experience supporting data platforms, data pipelines, or data-intensive workloads
- Knowledge of AWS services commonly used in data platforms, such as S3, RDS, Redshift, Glue, EMR, or related services
- Experience with Docker and container image management
- Experience with advanced observability and monitoring platforms
- Knowledge of SRE practices, including SLIs, SLOs, incident management, and reliability engineering
- Experience with AWS cost optimization and cloud resource management
- Experience implementing security and compliance best practices in AWS and Kubernetes environments
- AWS or Kubernetes-related certifications
- Experience working in large-scale, distributed, or multi-environment cloud platforms