We are looking for a Lead DevOps Engineer to partner with the client infrastructure team and raise the bar on CI/CD quality, security, and observability across three pods. You will evolve GitLab pipelines, consolidate infrastructure-as-code, establish delivery metrics, and help harden non-production environments.
Responsibilities
- Maintain GitLab CI/CD pipelines across all pods to keep build, test, and deploy flows stable and repeatable
- Implement and configure pipeline security scanning, including SAST/DAST and dependency vulnerability scanning
- Set up and sustain automated dependency update tooling (Renovate or Dependabot) across all repositories
- Drive infrastructure-as-code consolidation across pods to reduce configuration sprawl and improve maintainability
- Collaborate with the client infrastructure team to support the non-production environment access restriction initiative
- Configure ECS log drivers for products that currently lack centralized log collection
- Build and operate the DORA metrics infrastructure to increase visibility into delivery health
- Standardize observability instrumentation across pod services, covering logging, metrics, and distributed tracing
- Assist incident response by providing infrastructure-level diagnosis during the initial ownership period
- Embed within the client infrastructure team to align on standards and prevent shadow infrastructure
- Advise pod leads and engineers on deployment patterns, rollback procedures, and environment promotion
- Contribute to runbook documentation for infrastructure-level operational procedures
Requirements
- Proven DevOps or infrastructure engineering experience with 5+ years in similar roles
- Deep expertise in GitLab CI/CD pipeline setup, tuning, and ongoing maintenance
- Hands-on proficiency with AWS services, including ECS, ECR, Lambda, CloudFront, and S3
- Strong capability with Terraform and infrastructure-as-code practices
- Practical knowledge of Docker and container-based deployment approaches
- Working familiarity with Datadog, Snyk, and AWS Secrets Manager
- Solid understanding of observability practices, including logging, metrics, and distributed tracing
- Demonstrated background supporting incident response and performing infrastructure-level diagnosis
- Experience using AI tools for pipeline configuration generation, IaC authoring, or incident root cause analysis
- English proficiency at B2 (Upper-Intermediate) level or higher
Nice to have
- Familiarity with AI-assisted log analysis and anomaly detection in observability platforms
- Interest in automating repetitive DevOps tasks using LLM-based scripting or AI coding assistants