We are looking for a Senior Database Operations Engineer to join our team and execute a fleet-wide Aurora MySQL remediation and optimization program across a governed production estate. The Engineer owns the schema-change, indexing and cluster-isolation workstreams spanning 36 databases, delivering every change as an automated, reversible wave. Work is AWS-native throughout: Aurora clusters managed as Terraform-defined infrastructure, with Datadog and CloudWatch providing the validation signal for each promotion.
Responsibilities
- Nullify the stale JSON column across 36 databases in six controlled waves, validating each wave before promoting the next
- Build automated reconciliation and rollback tooling so every wave is independently verifiable and fully reversible
- Remove unused indexes to cut storage footprint and lift write throughput against a measured baseline
- Administer Aurora MySQL clusters end to end, covering patching, Blue/Green deployments and parameter group management
- Isolate two to three high-impact programs onto dedicated clusters to contain performance contention and blast radius
- Diagnose and tune query and index performance, evidencing the improvement in each optimization cycle
- Provision and evolve database infrastructure as code using Terraform and Terragrunt through GitHub CI/CD pipelines
- Publish and maintain database service definitions in the Backstage developer portal and service catalog
- Build Datadog and CloudWatch dashboards and alarms, including the post-patch validation checks that gate each wave
- Operate inside PCI-governed production access controls and evidence every change for audit
Requirements
- 3+ years of hands-on, relevant experience in a similar database operations role
- Practical experience working with AWS Aurora for managing and optimizing relational database clusters
- Solid background using Amazon Web Services to provision and manage cloud infrastructure
- Experience with Datadog for monitoring system performance and setting up alerting
- Strong MySQL knowledge, including hands-on database administration covering patching, tuning, and cluster management
- Experience using Terraform to provision and manage infrastructure as code
- Solid communication skills with working English fluency at B2 level or higher, enabling clear understanding of business requirements and their translation into technical architectures
Nice to have
- Experience with Amazon CloudWatch for tracking infrastructure metrics and configuring alarms
- Familiarity with Backstage for publishing and maintaining service catalog definitions
- Understanding of PCI (Payment Card Industry) compliance requirements within production environments
- Experience using Terragrunt to manage and orchestrate Terraform configurations