We are seeking a Senior Site Reliability Engineer (SRE) to oversee the reliability, performance, scalability, and security of the client's high-traffic, enterprise-scale Acquia (Drupal) web platform. This role is centered on managing and optimizing a sophisticated, multi-tier delivery architecture, with a primary and immediate focus on conducting comprehensive platform assessments of the current infrastructure, delivery pipelines, and caching configurations spanning Cloudflare, Varnish, and Memcached to identify bottlenecks, security gaps, and optimization opportunities.
Responsibilities
- Evaluate and document the existing infrastructure, including the integration between Cloudflare, Varnish, Memcached, and the Acquia hosting platform
- Analyze cache hit ratios (CHR) at the Cloudflare and Varnish layers, identify uncacheable components, and recommend changes to maximize edge and proxy caching
- Compare current configurations against Drupal, Acquia, and industry-standard SRE best practices, generating actionable roadmaps and recommendations for remediation
- Maintain, configure, and optimize the end-to-end caching pipeline (Cloudflare Edge → Varnish → Memcached → Drupal Backend) to minimize origin load
- Implement and manage robust cache invalidation strategies, such as Drupal Cache Tags, acquia_purge, and Cloudflare API purges, to ensure content freshness without risking origin stampedes
- Author and tune Varnish Configuration Language (VCL) files to manage custom headers, redirect logic, and bypass rules
- Set up and maintain monitoring, logging, and alerting dashboards across the entire stack, including Acquia logs, Varnish logs, Cloudflare Analytics, and APM tools like New Relic
- Participate in incident response to diagnose and resolve production bottlenecks, service disruptions, or latency spikes
- Run thorough Post-Incident Reviews (PIR) to prevent recurrences
Requirements
- 3+ years of experience auditing complex web architectures, generating technical assessment reports, and translating findings into prioritized technical backlogs
- Experience working within the Acquia Cloud ecosystem
- Understanding of Varnish caching, VCL scripting, and administrative tuning
- Skills in configuring and troubleshooting Memcached
- Knowledge of Cloudflare configuration
- English proficiency at B2 level or higher
Nice to have
- Background in hosting and tuning enterprise-grade Drupal sites
- Familiarity with Drupal internals for debugging, including how Drupal bootstraps, how settings.php overrides caching behavior, and how to debug PHP-FPM slow logs, memory exhaustion, and DB locks during assessments
- Proficiency in New Relic (native to Acquia) or Datadog to profile slow transactions, database query times, and external API call latencies during auditing
- Experience parsing and analyzing web logs, including Apache/Nginx logs on Acquia, Varnishncsa logs, and Cloudflare Logpush
- Familiarity with Git, Acquia Cloud Pipelines, or alternative tools like GitLab CI/CD, GitHub Actions, and deployment orchestration