Lead Database Operations Engineer

EPAM·Argentina, Colombia, Mexico, Chile·Удалённо·вчера

We're seeking a Lead Database Operations Engineer to become part of our team and drive a comprehensive Aurora MySQL remediation and optimization initiative spanning a tightly controlled production environment. This role takes ownership of schema modification, indexing, and cluster-isolation workstreams covering 36 databases, ensuring every change rolls out as an automated, fully reversible wave. All work is built natively on AWS: Aurora clusters function as Terraform-defined infrastructure, with Datadog and CloudWatch supplying the validation signals needed to greenlight each promotion

Responsibilities

  • Clear out the outdated JSON column spanning 36 databases across six managed waves, confirming each wave's success before advancing to the next
  • Develop automated reconciliation and rollback mechanisms ensuring every wave can be independently checked and completely undone if needed
  • Eliminate unused indexes to reduce storage consumption and boost write throughput relative to an established baseline
  • Take full ownership of Aurora MySQL cluster administration, including patch management, Blue/Green deployment execution, and parameter group configuration
  • Separate two to three critical programs onto their own dedicated clusters to limit performance interference and reduce potential impact radius
  • Analyze and improve query and index performance, documenting measurable gains achieved through each tuning cycle
  • Set up and refine infrastructure-as-code deployments using Terraform and Terragrunt within GitHub CI/CD pipeline workflows
  • Maintain and update database service listings within the Backstage developer portal and its service catalog
  • Construct dashboards and alarm configurations in Datadog and CloudWatch, incorporating post-patch validation checks that control each wave's progression
  • Function within PCI-regulated production access frameworks, documenting every change thoroughly to support audit requirements

Requirements

  • Five-plus years of practical, hands-on experience in a comparable database operations position
  • A minimum of one year spent leading and managing teams
  • Hands-on background working with AWS Aurora to manage and fine-tune relational database clusters
  • Strong track record using Amazon Web Services to set up and administer cloud-based infrastructure
  • Background using Datadog to track system performance metrics and configure alerting mechanisms
  • Deep MySQL expertise, including direct database administration experience covering patching, performance tuning, and cluster oversight
  • Background applying Terraform to build and manage infrastructure through code
  • Effective communication skills with working English fluency at B2 level or above, supporting clear grasp of business needs and their conversion into technical solutions

Nice to have

  • Background using Amazon CloudWatch to monitor infrastructure metrics and set up alarm thresholds
  • Exposure to Backstage for maintaining and publishing entries within a service catalog
  • Awareness of PCI (Payment Card Industry) compliance standards as applied to production systems
  • Background applying Terragrunt to coordinate and manage Terraform-based configurations

Похожие вакансии

Другие вакансии EPAM